Medical AIs can be queried to reveal whose data trained them, Nature paper finds
Discriminative medical AI models can be tricked into confirming whether a specific patient's records were part of their training data, German researchers report in a Nature paper published on 24 June. On individual patients, the attack succeeded almost every time.

The finding, first reported by The Register on 24 June, matters less for what it proves about model security in the abstract than for what it implies about the privacy promises that hospitals, insurers and AI vendors have been making to patients. If membership in a training set can be confirmed, the training set itself becomes a disclosure channel. That is the uncomfortable part.
Moritz Knolle chairs AI in Healthcare and Medicine at the Technical University of Munich and led the paper. He told The Register the exposure is not hypothetical. "You could imagine scenarios where membership in a training dataset reveals that someone has a dormant genetic condition such as Huntington's disease, depression, or attended a specific, specialised treatment clinic," he said. The scenario is not exotic. It is the ordinary business of clinical record keeping.
The attack class is not new. Membership inference attacks, or MIAs, have been studied in the machine learning literature for years, and the general shape of the technique is well understood. A model tends to be more confident about inputs it has seen before, so an attacker who can query the model and read its confidence scores can infer inclusion. What the Nature paper adds is scale and clinical specificity. The team tested seven medical AI datasets covering images, ECG records and general electronic health records. Individual patients could be identified with what the researchers describe as near-perfect attack success. That result is not a laboratory curiosity. It is a measurement of how these models behave when someone asks them the right question.
Why aggregate audits miss it
That result cuts against how these models are currently evaluated. Privacy audits typically measure attack success in aggregate across a dataset, which produces a reassuring average. The paper argues that this is the wrong unit of analysis. "The fact that MIAs can achieve near-perfect success rates for individual patients is not adequately captured by the standard evaluation protocol, which measures attack success in aggregate across records," the researchers write. Their conclusion is that reporting standards for AI privacy audits need to change.
A second finding sits awkwardly next to the first. Patients who are underrepresented in a training cohort are easier to identify than those who are not, because their data points stand out. Race, insurance status, sex, the imaging protocol used and certain disease statuses can all function as outliers. Knolle's summary of the pattern is blunt: "Generally speaking, privacy risks from MIAs become more severe as a model's training cohort becomes more specific."
The larger the dataset, the easier individual records are to expose, according to the paper. The researchers note that the magnitude of this change in patient-level risk in larger models was previously unknown. That is a counterintuitive result for anyone who assumes that scale brings anonymity. In medical imaging and genomics, it often does the opposite.
What an attacker actually needs
This is where the practical picture gets more complicated, and where the paper's own caveats matter. Running an MIA against a medical model requires the attacker to already hold some data belonging to the person they want to identify. Knolle confirmed as much. "To conduct a MIA an attacker needs access to a target data point," he told The Register. He added that the paper shows a full patient record is not required, contrary to prior assumptions. "In our paper we show that an attacker with partial access can still successfully conduct MIAs."
The attacker would simply need access to someone's blood test results, or part of these results.
That lowers the bar considerably. The attack works by feeding the partial record to the model, reading the confidence level, and inferring that the patient is in the training set. The datasets used in the study were anonymised, and Knolle points out that the attacker does not need to know who the data belongs to in order to run the query. The linkage between anonymised training data and a named individual has to come from somewhere else, but that somewhere else is often a breach.
Healthcare data breaches are frequent enough that this is not a theoretical concern. Knolle sketches the scenario plainly. An attacker gains unauthorised access to a general practitioner's database after a routine blood test, and from there can probe the model.
Asked what he wants the research to achieve, Knolle said he hopes the medical AI community will take privacy risks seriously and apply mitigation techniques where they are needed. The paper recommends differential privacy frameworks, which are designed to give mathematical guarantees that training data remains anonymous, and a reworked audit standard that reports individual-level risk rather than aggregate risk. A third option, which Knolle raised, is to compile training data so that underrepresented groups are better represented, reducing the outlier effect that makes them easy to identify.
He also cautioned against treating every successful MIA as a serious violation. "There are many situations where a successful MIA represents a small or negligible privacy violation," he noted, pointing to models trained on large, general populations where both healthy and diseased individuals are represented.
The regulatory gap
Set against the current regulatory picture, the paper lands in a space that rules have so far struggled to fill. Privacy audits for medical AI tend to focus on whether data was lawfully collected and whether identifiers were stripped, not on whether the model itself can be interrogated to reconstruct membership. Differential privacy is available as a technique, but it is not uniformly required. The paper's argument is essentially that the evaluation regime is measuring the wrong thing.
That gap is not unique to Europe. Regulators in several jurisdictions have been working through how to classify and supervise AI tools used in clinical settings, with proposals ranging from pre-market approval regimes to lighter-touch guidance for lower-risk applications. None of the frameworks in circulation, as far as the paper's authors are concerned, addresses the individual-level membership risk that their experiments demonstrate.
A parallel body of evidence suggests that the diagnostic performance of these systems is also weaker than marketing implies. Research published in JAMA Network Open in April, led by Harvard medical student Arya Rao and covered by The Register on 15 April, tested 21 off-the-shelf AI models on 29 standardised clinical vignettes. The models performed well when given a full portfolio of information and asked for a final diagnosis, with leading models correct 91 percent of the time. Early differential diagnosis, the stage where clinicians weigh competing possibilities, failed in more than 8 out of 10 cases.
"Every model we tested failed on the vast majority of cases," Rao told The Register. "That's the stage where uncertainty matters most, and it's where these systems are weakest." Dr Marc Succi, a radiologist at Massachusetts General Hospital and a coauthor, said the results suggest today's off-the-shelf models should not be trusted for patient-facing diagnostic reasoning without structured comprehensive human review.
Rao also noted that the stricter failure metric was not the whole story. Measured as raw accuracy, the proportion of cases where a model got the answer fully correct ranged from 63 to 78 percent. She described that as suggesting the models were often partially correct even when they failed under the stricter definition.
The two papers point in the same direction from different angles. One shows that the models can leak who was in their training data. The other shows that they are unreliable at the reasoning stage where clinical value is highest. Neither finding is fatal on its own, and neither rules out useful deployment in narrow, well-supervised settings.
What they do undermine is the assumption that current privacy audits and performance benchmarks together are sufficient to certify a medical AI as safe. The Nature paper's authors want audit standards rewritten to report individual-level risk. The JAMA team wants marketing claims about diagnostic agents to match the evidence. Both are requests for the same thing: evaluation that reflects how these systems fail, not how they succeed on average.
Sources
2- 01Medical diagnosis AIs can be tricked into telling whose data trained themEN
- 02LLMs fail in 8 out of 10 early differential diagnosis casesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.