MLA-C01 ML Model Development Practice Question
A team is training a large language model using PyTorch on SageMaker. They need to reduce training time. The model has 10 billion parameters. Which distributed training strategy should they use?
⚠ Common exam trap
MLA-C01 often tests the misconception that adding more GPUs via data parallelism solves all scaling problems, when in fact model size exceeding single-device memory requires model parallelism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model parallelism with SageMaker distributed
Model parallelism with SageMaker distributed is correct because a 10-billion-parameter model cannot fit into the memory of a single GPU, so the model itself must be partitioned across multiple GPUs/devices. SageMaker's distributed model parallelism library shards model layers and parameters across instances, enabling training of models that exceed single-device memory. This directly addresses the memory bottleneck that prevents scaling with data parallelism alone.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Data parallelism with Horovod
Why it's wrong here
Horovod replicates the full 10-billion-parameter model on every GPU, so per-device memory cannot hold it and training fails before speed matters. It suits smaller models fitting one device, where scaling across many GPUs cuts epoch time.
- ✗
Single GPU training
Why it's wrong here
Single-GPU training cannot hold a 10-billion-parameter model plus optimiser states and gradients within one device's memory, so training would fail or require impractical offloading. It is tempting because single-device training avoids communication overhead and suits small models that fit comfortably in one GPU's memory.
- ✗
Use a larger instance type without parallelism
Why it's wrong here
A single larger instance still holds one copy of the 10-billion-parameter model on its GPUs, capping memory and throughput. It suits models fitting comfortably on one node, where avoiding inter-node communication overhead genuinely shortens training.
- ✓
Model parallelism with SageMaker distributed
Why this is correct
SageMaker distributed model parallelism partitions the 10-billion-parameter model across GPUs, holding each shard separately so the full weight set need not fit on one device. This suits the stem's large language model, where data parallelism would replicate all parameters per GPU and exhaust memory.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
2 more ways this is tested on MLA-C01
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A team is training a large language model and needs to split the model layers across multiple GPUs due to memory constraints. Which distributed training strategy should they use?
medium- A.Data parallelism
- B.Hyperparameter tuning
- C.Autopilot
- ✓ D.Model parallelism
Why D: Model parallelism splits the model's layers (or tensors) across multiple GPUs so that each GPU holds only a portion of the model's parameters, which is required when the model is too large to fit in a single GPU's memory. Data parallelism, by contrast, replicates the full model on every GPU and only partitions the training data, so it does not solve a memory-constraint problem.
Variation 2. A team is training a large language model on SageMaker using PyTorch with data parallelism. The model is too large to fit on a single GPU. Which distributed training strategy should they use to split the model across multiple GPUs?
medium- ✓ A.Model parallelism
- B.Tensor parallelism
- C.Data parallelism
- D.Pipeline parallelism
Why A: Model parallelism splits the model itself across devices, which is necessary when the model is too large for one GPU. SageMaker's model parallelism library supports this.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.