hardMultiple Choice
MLA-C01 Practice Question: A team is using SageMaker to run a large-scale…
A team is using SageMaker to run a large-scale distributed training job for a language model. They are using SageMaker's Pipe mode to stream data from S3 to reduce IO. They observe that the training throughput is lower than expected, and the CPU utilization is high while GPU utilization is low. The training script uses PyTorch's DataLoader with num_workers=0. The data preprocessing is minimal. Which change is most likely to improve GPU utilization?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of data loading workers (num_workers).
With num_workers=0, PyTorch's DataLoader loads data in the main training process, creating a CPU bottleneck that keeps GPUs idle. Increasing num_workers parallelizes data loading across multiple subprocesses, which reduces CPU strain and feeds data faster to GPUs, improving throughput. Adding more GPUs (Option C) or vCPUs (Option B) does not address the root cause, and switching to File mode (Option D) would increase I/O overhead, worsening performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Increase the number of data loading workers (num_workers).
Why this is correct
With num_workers=0, PyTorch loads and preprocesses batches synchronously on the main process, starving the GPU while CPUs work. Raising num_workers enables parallel background data loading, keeping the GPU fed and lifting utilisation during distributed training.
- ✗
Use a larger instance with more vCPUs.
Why it's wrong here
More vCPUs do not raise GPU utilisation because num_workers=0 confines data loading and preprocessing to the main process on a single core; the extra cores stay idle. Larger CPU instances suit CPU-bound preprocessing that already runs across multiple DataLoader worker processes.
- ✗
Increase the number of GPUs per instance.
Why it's wrong here
Adding GPUs cannot help while the input pipeline starves them: with num_workers=0, the main process loads and preprocesses every batch serially, so extra devices simply idle alongside the existing ones. Scaling GPU count suits compute-bound jobs where kernels, not data loading, limit step time.
- ✗
Switch from Pipe mode to File mode.
Why it's wrong here
File mode still feeds the same DataLoader, so with num_workers=0 the single process remains the bottleneck and GPU starvation persists. It is tempting because File mode suits random-access workloads, and would be correct if the script needed non-sequential reads rather than higher preprocessing parallelism.
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.