A healthcare company must train a model on sensitive patient data while complying with privacy regulations. They want to add noise to the training process to prevent re-identification. Which technique should they implement?
Differential privacy injects calibrated statistical noise into training or query outputs, bounding any single patient's influence on the model. This mathematically limits re-identification risk while preserving aggregate utility, satisfying privacy regulations when training on sensitive patient records.
Why this answer
Differential privacy is the correct technique because it adds calibrated noise to the training process (e.g., via gradient clipping and noise injection in stochastic gradient descent) to ensure that the model's outputs do not reveal whether any individual's data was included in the training set. This provides a formal mathematical guarantee (ε-differential privacy) that limits the risk of re-identification, which is essential for complying with privacy regulations like HIPAA or GDPR when training on sensitive patient data.
Exam trap
AWS often tests the misconception that federated learning alone provides privacy guarantees, but the trap here is that federated learning only addresses data locality, not re-identification resistance, which requires a formal privacy technique like differential privacy.
How to eliminate wrong answers
Option B (k-anonymity) is wrong because it is a data anonymization technique applied to static datasets (e.g., generalizing quasi-identifiers in a table) rather than a training process technique; it does not add noise during model training and can be vulnerable to attacks like homogeneity or background knowledge attacks. Option C (Federated learning) is wrong because it is a distributed training approach that keeps data on local devices but does not inherently add noise to prevent re-identification; without differential privacy, model updates can still leak sensitive information. Option D (Homomorphic encryption) is wrong because it allows computation on encrypted data but does not add noise to the training process; it protects data in transit or at rest but does not prevent re-identification from model outputs.