Medical AI privacy audits miss what matters most: individual patients
Diagnostic AI models used in medicine can be tricked into revealing whether a specific patient's data was part of their training set, and German researchers say the standard way those models are audited for privacy fails to catch it.

German researchers published a Nature paper on Wednesday. Discriminative medical AI models, the kind used to classify data and make predictions about new inputs, are "particularly susceptible" to membership inference attacks (MIAs), they report. The Register reported on the findings on 24 June.
An MIA queries a model to work out whether a particular datapoint was included in its training set. For medical AI, that means a patient whose records helped train the model could be exposed, with details of their medical history and diagnoses leaking as a result. The team analysed seven medical AI datasets covering images, ECG records and general electronic health records. Individual patients targeted by such attacks could be identified with what the researchers describe as "near-perfect attack success."
Aggregate metrics hide the problem
The researchers argue this success rate is not captured by the way medical AI models are currently evaluated for safety. "The fact that MIAs can achieve near-perfect success rates for individual patients is not adequately captured by the standard evaluation protocol, which measures attack success in aggregate across records," they wrote. On that basis they conclude that reporting standards for AI privacy audits need to change. Underrepresented patients are easier to identify than others. Race, insurance status, sex, the protocol used to conduct medical imaging, and certain disease statuses can all function as outliers, according to the paper. Those outliers make membership inference easier.
"Generally speaking, privacy risks from MIAs become more severe as a model's training cohort becomes more specific," Technical University of Munich AI in Healthcare and Medicine chair and paper lead author Moritz Knolle told The Register by email.
Knolle gave examples of what that specificity could expose. "You could imagine ... scenarios where membership in a training dataset reveals that someone has a dormant genetic condition such as Huntington's disease, depression, or attended a specific, specialised treatment clinic," he said. Larger datasets do not fix the problem. According to The Register's write-up of the paper, the larger the dataset, the easier it is to expose records, and "the magnitude of this change in patient-level risk was previously unknown" in larger models.
What an attacker actually needs
An MIA is not a remote, no-input attack. Knolle confirmed that to conduct one, an attacker needs access to a target data point. But the research also undercuts an assumption that previously made these attacks look impractical: access to a full patient record is not required. "In our paper we show that an attacker with partial access can still successfully conduct MIAs," Knolle told The Register.
The attack itself exploits model confidence. Medical AIs tend to be more certain of their predictions when the input data already formed part of their training set. An attacker feeds the model patient data they have obtained, checks the confidence level, and infers from it that the patient is in the training data.
"An attacker conducting a MIA does not need to know who the data belongs to that they are trying to conduct the MIA with," Knolle said. "In fact, all the dataset we use in our study were anonymized." The datasets were anonymised. The target data was not. The paper reports that the MIA attacks were largely error-free at the individual patient level, which means confidence levels are an accurate signal of whether a given patient's data sits in a training set. In practice, Knolle said, "The attacker would simply need access to someone's blood test results, or part of these results" to infer inclusion. Getting hold of that data is the remaining hurdle, and not a large one in Knolle's view. "Given that medical data is not always securely stored it is not unthinkable that an attacker could get access, for example, by gaining unauthorized access to the database of your general practitioner after they performed a routine blood test," he said.
Recommendations and a caveat
The researchers make several recommendations. One is the use of differential privacy frameworks, designed to mathematically guarantee that training data stays anonymous. Another is reform of privacy audit standards so they consider individual-level risk rather than aggregate privacy risk alone. A third option, Knolle said, is to compile medical AI training data so that underrepresented groups are better represented. Knolle also offered a limit on how alarming any single MIA should be. "There are many situations where a successful MIA represents a small or negligible privacy violation," he noted, pointing to AI models trained on large, general populations in which both healthy and diseased individuals are represented. Asked what he hopes the research achieves, Knolle said he wants the medical AI community to take privacy risks seriously and to use risk mitigation techniques where they are necessary.
The other failure mode
Privacy is not the only documented weakness in medical AI. In April, The Register reported on a study published in JAMA Network Open that examined 21 leading off-the-shelf AI models across 29 standardised clinical vignettes. The models did well at final diagnosis when given a full portfolio of medical information, with leading models correct 91 percent of the time. Early differential diagnosis was a different story. The study found a failure rate above 80 percent at that stage. "Every model we tested failed on the vast majority of cases," lead author Arya Rao, a Harvard medical student, told The Register. "That's the stage where uncertainty matters most, and it's where these systems are weakest."
Rao added that the strict failure metric was not the only way to read the data. Raw accuracy, as a proportion of cases correct, ranged from 63 to 78 percent, which she said suggests models were often partially correct even when they failed to produce a fully correct differential. Dr. Marc Succi, a Massachusetts General Hospital radiologist and coauthor, said the higher success rate at final diagnosis should not be reassuring. "Real clinical reasoning starts earlier, when ambiguity is highest, and that is exactly where they remain weakest," he told The Register, adding that a wrong differential can lead to delays in care, unnecessary procedures with complications and high costs.
Both papers point at the same gap between how medical AI is measured and how it is used. One measures privacy in aggregate while individual patients remain identifiable. The other scores final answers while the earlier, more uncertain stage of reasoning stays unreliable.
Sources
2- 01Medical diagnosis AIs can be tricked into telling whose data trained themEN
- 02LLMs fail in 8 out of 10 early differential diagnosis casesEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.