MLA-C01 ML Model Development Practice Question
A data scientist is using SageMaker to train a custom PyTorch model for image classification. They want to use SageMaker Debugger to detect training issues. Which TWO built-in rules are most relevant for detecting common training problems? (Select TWO.)
⚠ Common exam trap
The trap is selecting plausible-sounding but non-existent Debugger rules — candidates must know the actual built-in rule names (Overfit, ExplodingGradients, VanishingGradient, etc.) and not confuse them with evaluation metrics or data quality tools.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Overfit
Option B (Overfit) is correct because SageMaker Debugger's built-in Overfit rule monitors the gap between training and validation loss across steps and raises an issue when validation loss stops improving while training loss keeps decreasing, which is the classic signature of overfitting in a PyTorch image-classification job. Option C (ExplodingGradients) is correct because the ExplodingGradients rule inspects the gradients tensor emitted by the framework and flags abnormally large gradient values or spikes, which cause unstable or diverging training and are a common problem in deep networks. Option A (DataDistribution) is not a built-in Debugger rule for detecting training problems; it relates to SageMaker Clarify/Model Monitor data and bias analysis rather than Debugger's training-issue rule set. Option D (ImageQuality) is not a SageMaker Debugger built-in rule; image-quality checks would be a custom preprocessing concern, not a Debugger rule. Option E (ConfusionMatrix) is not a Debugger training-issue rule; confusion matrices are evaluation artifacts computed after training (for example with SageMaker Clarify or custom code), not a built-in Debugger rule for detecting training problems.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
DataDistribution
Why it's wrong here
DataDistribution is a valid SageMaker Debugger built-in rule, so it does not fail this scenario; it detects skew between training data shards across instances. It is relevant to distributed training problems, but the question asks for rules detecting common training issues such as vanishing gradients or overfitting.
- ✓
Overfit
Why this is correct
The Overfit rule in SageMaker Debugger monitors the validation loss relative to the training loss; if validation loss begins to increase while training loss continues to decrease, the rule emits a warning. This directly addresses the image classification scenario, where a custom PyTorch model can easily memorise training data rather than generalising, satisfying the stem’s requirement to detect common training problems.
- ✓
ExplodingGradients
Why this is correct
ExplodingGradients monitors gradient magnitudes during backpropagation and raises an alert when they exceed a threshold, flagging unstable training such as diverging loss or NaN weights. This directly targets the vanishing/exploding gradient problems common in deep PyTorch image classifiers.
- ✗
ImageQuality
Why it's wrong here
ImageQuality is not a SageMaker Debugger built-in rule; Debugger monitors tensors and gradients during training, not input image fidelity. It would be tempting when auditing dataset quality, but that belongs to SageMaker Clarify or Data Wrangler, not Debugger's training-issue rules.
- ✗
ConfusionMatrix
Why it's wrong here
ConfusionMatrix is a SageMaker Clarify bias and explainability metric, not a Debugger training rule; Debugger inspects tensors, gradients and losses during training. It tempts when evaluating classification accuracy post-training, but that is model evaluation, not detecting training issues.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on MLA-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. An ML engineer is debugging a training job that is consistently failing due to an out-of-memory error. The engineer is using SageMaker's built-in XGBoost algorithm. Which Debugger rule can help identify the issue?
medium- ✓ A.Exploding gradients
- B.Overfit
- C.Dead relu
- D.OOM rule
Why A: The 'Exploding gradients' rule detects when gradients become too large, which is a common cause of training instability but not necessarily OOM. The 'Overfit' rule detects overfitting. The 'Dead relu' rule is for ReLU activation. None of these directly address OOM. However, Debugger does not have a specific OOM rule; instead, the engineer should monitor memory utilization via CloudWatch or adjust instance type. Among the options, 'Exploding gradients' is the most relevant because large gradients can lead to memory spikes.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.