A healthcare AI system uses patient data to predict disease risk. To comply with HIPAA and reduce the risk of re-identification, which technique should be applied to the training data before model development?
Trap 1: Pseudonymisation by replacing patient names with random IDs
Replacing names with random IDs leaves other identifiers such as dates of birth, postcodes, and rare diagnoses intact, so records remain re-identifiable; HIPAA's de-identification standard requires removal or transformation of all such identifiers. Pseudonymisation suits internal linkage and audit trails, not releasing data for model training.
Trap 2: Data augmentation to create synthetic samples
Augmentation generates additional synthetic records from existing ones, which does not remove identifiers or break the link between a record and the individual, so re-identification risk is unchanged. It tempts because synthetic data can reduce exposure, but that requires fully generative replacement, not augmentation of real patient records.
Trap 3: Data minimisation by removing all features except age and gender
Stripping all features except age and gender destroys the clinical signal needed to predict disease risk, so the model cannot meet its purpose; HIPAA permits de-identification while retaining relevant attributes. It tempts because minimisation is a genuine privacy principle, but it applies to collecting only necessary fields, not discarding diagnostic data.
- A
Pseudonymisation by replacing patient names with random IDs
Why it fails: Replacing names with random IDs leaves other identifiers such as dates of birth, postcodes, and rare diagnoses intact, so records remain re-identifiable; HIPAA's de-identification standard requires removal or transformation of all such identifiers. Pseudonymisation suits internal linkage and audit trails, not releasing data for model training.
- B
Data augmentation to create synthetic samples
Why it fails: Augmentation generates additional synthetic records from existing ones, which does not remove identifiers or break the link between a record and the individual, so re-identification risk is unchanged. It tempts because synthetic data can reduce exposure, but that requires fully generative replacement, not augmentation of real patient records.
- C
Differential privacy with a carefully chosen epsilon
Differential privacy adds controlled noise to protect individual records, meeting HIPAA's de-identification standards with formal guarantees.
- D
Data minimisation by removing all features except age and gender
Why it fails: Stripping all features except age and gender destroys the clinical signal needed to predict disease risk, so the model cannot meet its purpose; HIPAA permits de-identification while retaining relevant attributes. It tempts because minimisation is a genuine privacy principle, but it applies to collecting only necessary fields, not discarding diagnostic data.