Which TWO of the following are effective defenses against adversarial examples in AI systems?
Adversarial training augments the training set with perturbed examples labelled correctly, forcing the model to learn robust decision boundaries. This directly reduces sensitivity to the small input perturbations that adversarial attacks exploit, hardening the classifier against them.
Why this answer
Adversarial training (option A) is a correct defense because it augments the training set with adversarial examples generated by attacks such as FGSM or PGD, so the model learns to classify perturbed inputs correctly and its decision boundary becomes more robust. Input sanitization and feature squeezing (option D) are also correct because they preprocess inputs to remove or reduce the adversarial perturbation — for example, by quantizing pixel values, spatial smoothing, or reducing color bit depth — which shrinks the attacker's effective search space and can neutralize small perturbations. Option B is not marked correct because model ensembles with majority voting can sometimes improve robustness but are not a reliable standalone defense; attackers can craft transferable or ensemble-aware adversarial examples that fool all members.
Option C is wrong because increasing sensitivity to input changes is the opposite of what is wanted — it makes the model easier to fool with tiny perturbations. Option E is wrong because pruning reduces model complexity for efficiency and may slightly alter robustness, but it is not an established defense against adversarial examples and can even degrade robustness.
Exam trap
The AI0-001 exam often tests the misconception that ensemble methods or model simplification inherently improve adversarial robustness, when in fact they do not address the fundamental mechanism of adversarial perturbations and may even weaken defenses.