hardMultiple Select
MLA-C01 Practice Question: A data scientist is training a large transformer…
A data scientist is training a large transformer model using SageMaker's model parallelism library. The training job is failing with an out-of-memory (OOM) error. Which two actions can help resolve the OOM error? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse pipeline parallelism with tensor parallelism, assuming decreasing pipeline degree reduces memory, when in fact it increases per-GPU memory load due to fewer stages.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the sequence length
Option A is correct because reducing the sequence length directly lowers the activation memory footprint of a transformer, since attention and intermediate activations scale with sequence length, making it a standard remedy for OOM during SageMaker model-parallel training. Option B is correct because activation checkpointing (gradient checkpointing) recomputes activations during the backward pass instead of storing all of them, substantially reducing memory usage at the cost of some extra compute. Option C is incorrect because increasing the batch size per GPU raises memory consumption and would worsen the OOM. Option D is incorrect because switching to a smaller instance type reduces available GPU memory, making OOM more likely. Option E is incorrect because decreasing the pipeline parallelism degree spreads the model across fewer stages, increasing per-GPU memory pressure rather than relieving it.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Reduce the sequence length
Why this is correct
Reducing sequence length shrinks activation memory, which scales linearly with token count, directly relieving the per-device memory pressure causing the OOM. Since model parallelism shards parameters but not activations, this addresses the constraint the stem identifies without altering the sharding configuration or requiring additional instances.
- ✓
Enable activation checkpointing
Why this is correct
Activation checkpointing stores only selected activations during the forward pass and recomputes the remainder during backpropagation, trading extra compute for a substantially smaller memory footprint. This directly relieves the OOM constraint by reducing peak activation memory held across transformer layers, allowing the model-parallel job to fit within device memory.
- ✗
Increase the batch size per GPU
Why it's wrong here
Larger batches raise activation memory per GPU, worsening the OOM rather than relieving it. It is tempting because increasing batch size improves throughput once memory headroom exists, so it becomes viable only after sharding or checkpointing has freed capacity.
- ✗
Switch to a smaller instance type
Why it's wrong here
Switching to a smaller instance type reduces GPU memory, worsening the OOM error rather than resolving it. It is tempting because smaller instances cut cost for lightweight workloads, and would be the right choice when a job is over-provisioned and underutilising capacity, but model parallelism requires sufficient aggregate memory across devices.
- ✗
Decrease the pipeline parallelism degree
Why it's wrong here
Lowering pipeline parallelism reduces the number of model layers split across stages, so each stage holds more layers and its activations and parameters grow, worsening memory pressure. It is tempting because fewer pipeline stages can cut cross-stage communication, but memory per device rises, not falls.
Visual reference
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.