Medical AI models leak which patients trained them, Nature study finds
German researchers have shown that AI models used to help diagnose medical conditions can be queried to reveal whether a specific patient's data was part of their training set. At the individual level, they report near-perfect success.

The finding appeared in a Nature paper on Wednesday 24 June, and The Register reported it the same day. It covers seven medical AI datasets built from images, ECG records and general electronic health records. The team behind the work says current privacy audits miss the risk because they measure attack success in aggregate rather than patient by patient.
The attack class is not new. What is new is how well it works on medical models.
Membership inference attacks (MIAs) query a model to work out whether a particular datapoint was in its training set. Discriminative models, the kind used to classify data and make predictions about new inputs, are especially easy to probe this way, according to the researchers. In their analysis, individual patients targeted by such attacks could be identified with what the paper calls near-perfect attack success. That matters because a confirmed training-set membership is itself a disclosure. If a model was trained on a cohort from a specialised clinic, proving that a person's record sits in that cohort can reveal a diagnosis they never chose to share.
Underrepresented patients are easier to pick out
The paper also found that patients who are outliers in a training set are easier to identify than those who are not. Race, insurance status, sex, the imaging protocol used and certain disease statuses can all function as outliers, according to the researchers.
Moritz Knolle leads the AI in Healthcare and Medicine chair at the Technical University of Munich and is lead author on the paper. He put the trade-off plainly in an email exchange with The Register. "Generally speaking, privacy risks from MIAs become more severe as a model's training cohort becomes more specific," he said. He offered an example: "You could imagine ... scenarios where membership in a training dataset reveals that someone has a dormant genetic condition such as Huntington's disease, depression, or attended a specific, specialised treatment clinic."
The researchers report another effect running the wrong way. Larger datasets make records easier to expose, not harder. The paper states that the size of this shift in patient-level risk had not previously been measured for larger models.
What an attacker actually needs
None of this means a stranger can point at a model and pull out names. An MIA requires the attacker to already hold a data point they want to test. Knolle confirmed as much, while noting the paper narrows how much data that takes.
"To conduct a MIA an attacker needs access to a target data point," he told The Register. "In our paper we show that an attacker with partial access can still successfully conduct MIAs."
The mechanism is confidence. Medical AIs tend to be more certain when the input was in their training set, so an attacker feeds in obtained patient data, reads the confidence level and infers inclusion. Knolle said the attacker does not need to know who the data belongs to. All the datasets used in the study were anonymised, he added, which does not protect against this kind of inference. "The attacker would simply need access to someone's blood test results, or part of these results" to infer inclusion, Knolle said. Asked where that data comes from, he pointed to routine exposure rather than exotic intrusion: "Given that medical data is not always securely stored it is not unthinkable that an attacker could get access, for example, by gaining unauthorized access to the database of your general practitioner after they performed a routine blood test."
Recommendations, and a caveat
The team's recommendations are not exotic. They want differential privacy frameworks used where the risk justifies it, privacy audit standards rewritten to assess individual-level rather than aggregate risk, and training data compiled so that underrepresented groups are better represented rather than left as outliers.
Knolle framed his goal narrowly. "I hope that the medical AI community will start to take privacy risks seriously and that risk mitigation techniques are used in situations where they are necessary," he said. He also pushed back on treating every successful attack as a scandal. "There are many situations where a successful MIA represents a small or negligible privacy violation," he noted, describing models trained on large, general populations in which both healthy and diseased individuals appear.
The regulatory context the paper lands in is crowded but thin on specifics. The Register's report notes that reporting standards for AI privacy audits need to change, which is a call on auditors and journals as much as on vendors. Nothing in the paper proposes a new certification regime, and no regulator is named as having adopted its recommendations.
There is a second, separate body of work pointing the same way from a different angle: whether these systems should be making diagnostic calls at all. A study led by Harvard medical student Arya Rao was published in JAMA Network Open and covered by The Register on 15 April. It tested 21 off-the-shelf AI models against 29 standardised clinical vignettes. The models were correct 91 percent of the time when given a full portfolio of information and asked for a final diagnosis.
Early differential diagnosis, the stage where clinicians rule conditions in and out while weighing uncertainty, failed in more than 8 out of 10 cases under the paper's strict definition. "Every model we tested failed on the vast majority of cases," Rao told The Register. "That's the stage where uncertainty matters most, and it's where these systems are weakest."
Rao also cautioned against reading the failure rate as total collapse. Measured as raw accuracy per case, the models scored between 63 and 78 percent, which she said suggests they were often partially right. Coauthor Dr Marc Succi, a radiologist at Massachusetts General Hospital, said higher final-diagnosis scores should not reassure anyone. "Real clinical reasoning starts earlier, when ambiguity is highest, and that is exactly where they remain weakest," he told The Register.
Put the two papers side by side and the regulatory question sharpens. A model that can leak who trained it and still misses most early differentials is being pitched for frontline triage. The audits used to clear it measure neither risk at the level where it bites.
Sources
2- 01Medical diagnosis AIs can be tricked into telling whose data trained themEN
- 02LLMs fail in 8 out of 10 early differential diagnosis casesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.