A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)
The Overfit rule in SageMaker Debugger monitors the validation loss relative to the training loss; if validation loss begins to increase while training loss continues to decrease, the rule emits a warning. This directly addresses the image classification scenario, where a custom PyTorch model can easily memorise training data rather than generalising, satisfying the stem’s requirement to detect common training problems.
Why this answer
Option B (Overfit) is correct because SageMaker Debugger's built-in Overfit rule monitors the gap between training and validation loss across steps and raises an issue when validation loss stops improving while training loss keeps decreasing, which is the classic signature of overfitting in a PyTorch image-classification job. Option C (ExplodingGradients) is correct because the ExplodingGradients rule inspects the gradients tensor emitted by the framework and flags abnormally large gradient values or spikes, which cause unstable or diverging training and are a common problem in deep networks. Option A (DataDistribution) is not a built-in Debugger rule for detecting training problems; it relates to SageMaker Clarify/Model Monitor data and bias analysis rather than Debugger's training-issue rule set.
Option D (ImageQuality) is not a SageMaker Debugger built-in rule; image-quality checks would be a custom preprocessing concern, not a Debugger rule. Option E (ConfusionMatrix) is not a Debugger training-issue rule; confusion matrices are evaluation artifacts computed after training (for example with SageMaker Clarify or custom code), not a built-in Debugger rule for detecting training problems.
Exam trap
The trap is selecting plausible-sounding but non-existent Debugger rules — candidates must know the actual built-in rule names (Overfit, ExplodingGradients, VanishingGradient, etc.) and not confuse them with evaluation metrics or data quality tools.