EU AI Act Compliance and Medical Diagnosis AI: What the Rules Actually Say
Medical diagnosis AI is being sold faster than regulators can audit it. Three recent studies point to the core problem: the models are hard to check, hard to secure and hard to quantify.

The regulatory question is not whether AI can diagnose. It is whether anyone can prove how well it does so. Two papers published months apart show how far the gap has opened.
On 24 June 2026, researchers at the Technical University of Munich reported in Nature that medical AI models used for diagnosis are vulnerable to membership inference attacks (MIAs). An MIA is a query that tries to work out whether a specific patient's record was part of the training set. The Register covered the paper the same day. The team tested seven medical AI datasets covering images, ECG records and general electronic health records. Individual patients, the authors found, can be identified with near-perfect attack success.
That is a privacy problem. It is also a measurement problem. The standard evaluation protocol measures attack success in aggregate across records. As the researchers point out, that approach does not adequately capture near-perfect success rates for individual patients. An audit that reports an average can miss the individual case entirely.
What the rules require, and what they miss
The EU AI Act is the framework most often cited in this debate, and the dossier here is thin on its operational detail. What the sources do show is a gap between the compliance language vendors use and the security research published in the same period.
A GitHub project, srta-ai-accountability, describes itself as the first system to address why-questions in explainable AI with 94% EU AI Act compliance. It is explicitly labelled a research prototype and a conceptual architecture. That is the state of play for at least some accountability tooling: a claim of 94% compliance, published on a code repository, with no regulatory body attached to the number.
The Munich paper points in a different direction. Its authors want reporting standards for AI privacy audits to change, so that individual-level risk counts, not just aggregate risk. They also recommend differential privacy frameworks, which are designed to guarantee mathematically that training data stays anonymous. And they suggest that training data could be compiled so underrepresented groups are better represented.
Lead author Moritz Knolle told The Register that privacy risks from MIAs become more severe as a model's training cohort becomes more specific. He gave the example of membership in a training dataset revealing that someone has a dormant genetic condition such as Huntington's disease, depression, or attended a specific specialised treatment clinic.
The accuracy claims regulators will have to weigh
Diagnostic performance is the other half of the regulatory problem, and the numbers here are unusually messy.
Microsoft announced its AI Diagnostic Orchestrator (MAI-DxO) on 30 June 2025. GeekWire reported that the tool beat 21 experienced physicians from the US and UK on complex cases from the New England Journal of Medicine, scoring 85.5% against the doctors' 20%. WIRED, covering the same work on 1 July 2025, put the figure at 80% versus 20% and said the system cut costs by 20% by selecting cheaper tests. The benchmark was built from 304 recent NEJM cases.
Those figures carry a caveat that appears in both accounts. Doctors in the study were not allowed to consult colleagues or use outside resources, which is not how clinicians work. MIT scientist David Sontag told WIRED that Microsoft's findings should be treated with caution for exactly that reason, and that it remains to be seen whether the system would cut costs in practice. Microsoft itself said the tool would need clinical testing and regulatory approval before deployment.
Then there is the failure case. On 15 April 2026, The Register reported on a study in JAMA Network Open led by Harvard medical student Arya Rao. It tested 21 leading off-the-shelf AI models against 29 standardised clinical vignettes. The models were correct 91% of the time when asked for a final diagnosis with full information. But in early differential diagnosis, the stage where clinicians rule conditions in and out, they failed in more than 8 out of 10 cases.
Every model we tested failed on the vast majority of cases, Rao told The Register. That is the stage where uncertainty matters most, and it is where these systems are weakest.
Massachusetts General Hospital radiologist Dr Marc Succi, a coauthor, told the same publication that today's off-the-shelf models should not be trusted for patient-facing diagnostic reasoning without structured comprehensive human review. He also said higher success rates in final diagnosis should not be reassuring, because real clinical reasoning starts earlier, when ambiguity is highest.
What a regulator would have to test
Put the three findings side by side and the audit problem becomes concrete.
- Privacy: individual patients in a training set can be identified with high success, and underrepresented groups are easier to identify than others. Aggregate reporting hides this.
- Performance: the same class of model can score 91% on final diagnosis and fail more than 80% of early differential cases.
- Benchmark design: Microsoft's headline result came from a test where human doctors worked without the resources they normally use.
None of this means medical AI cannot be regulated. It means the current instruments measure the wrong unit. An aggregate privacy score, a final-answer accuracy rate and a benchmark run under artificial constraints are all easy to publish and hard to falsify.
Knolle said he hopes the medical AI community will start to take privacy risks seriously and that risk mitigation techniques are used where they are necessary. He also noted that there are many situations where a successful MIA represents a small or negligible privacy violation, such as models trained on large, general populations.
The unresolved question is who decides which situation a given model is in. Under the EU AI Act, that decision is supposed to sit with the provider and the auditor, not with the paper's authors and not with a repository that grades itself at 94%.
Sources
5- 01Medical diagnosis AIs can be tricked into telling whose data trained themEN
- 02Microsoft AI tool outperforms doctors in diagnosing complex medical casesEN
- 03Microsoft Says New AI Diagnosed Patients 4 Times More Accurately Than DoctorsEN
- 04LLMs fail in 8 out of 10 early differential diagnosis casesEN
- 05srta-ai-accountability: AI Accountability FrameworkEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.