MLA-C01 ML Model Development Practice Question
A machine learning engineer is using Amazon SageMaker Debugger to monitor a training job for a deep neural network. They receive a rule alert indicating 'exploding gradients'. Which action should they take to address this issue?
⚠ Common exam trap
MLA-C01 often tests the confusion between exploding and vanishing gradients, where candidates might think increasing learning rate or adding layers helps, but the correct action is to reduce the learning rate or apply gradient clipping.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the learning rate
Exploding gradients occur when large error gradients accumulate during backpropagation, causing unstable updates and divergence. Reducing the learning rate directly scales down the parameter update step (Δθ = -η∇J), preventing the weights from overshooting and stabilizing training. This is the standard first-line remedy for exploding gradients in deep networks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a smaller batch size
Why it's wrong here
Reducing batch size changes gradient averaging but does not rescale the exploding magnitudes, so divergence continues. It is tempting because smaller batches are commonly used to improve generalisation or fit memory limits, and that would be the right adjustment when the problem is memory exhaustion rather than exploding gradients.
- ✓
Reduce the learning rate
Why this is correct
Exploding gradients arise when large updates compound across layers, so lowering the learning rate shrinks each step and restores stability. This directly addresses the alert SageMaker Debugger raised, though gradient clipping is an alternative remedy.
- ✗
Increase the number of layers to absorb gradients
Why it's wrong here
Adding layers increases depth and multiplies the gradient products flowing backward, intensifying the explosion rather than absorbing it. It is tempting because extra layers raise model capacity, and that would be the correct move when the network underfits and needs greater representational power.
- ✗
Increase the learning rate
Why it's wrong here
Raising the learning rate amplifies each update, worsening the divergence that exploding gradients already cause. It is tempting because a higher rate speeds convergence when training is slow, and it would be appropriate when gradients are vanishing or the model is underfitting and needs faster progress.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.