A retail bank's fraud model was trained on customer transaction data that included account holders in the EU. An internal audit finds the training pipeline copied raw transaction records, including names and card numbers, into an unencrypted research bucket for model retraining. Which action best aligns the remediation with data-protection obligations for that pipeline?
The violation is storing identifiable personal data outside its lawful purpose and controls. Pseudonymizing or tokenizing before the data reaches the research environment, plus defined retention limits, reduces identifiability while preserving the statistical signal the fraud model needs. Deleting the exposed copy ends the ongoing exposure, and the documented retention schedule satisfies accountability and minimization expectations for the pipeline.
Why this answer
The finding combines excessive identifiability with a purpose and retention failure. Removing the exposed copy stops ongoing risk, while pseudonymization or tokenization plus documented retention limits lets the fraud model keep learning from transaction behavior without carrying direct identifiers into a research environment. Encryption, monitoring, and new consent each address part of the problem but leave the core data-minimization defect unresolved.
Exam trap
The trap here is treating encryption of the research bucket as a complete privacy fix, when the governing issue is that identifiable data was copied outside its original purpose and kept without a retention limit.