mediumMultiple Select
PMLE Practice Question: A data scientist needs to scale a prototype deep…
A data scientist needs to scale a prototype deep learning model to train on a massive dataset using multiple GPUs. Which three strategies are essential for efficient distributed training? (Select THREE)
⚠ Common exam trap
PMLE often tests the trade-off between synchronous and asynchronous updates, and candidates may incorrectly choose asynchronous as 'more efficient' when synchronous is preferred for convergence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement data parallelism.
Option B is correct because data parallelism is the foundational strategy for multi-GPU training: the model is replicated on each GPU and each replica processes a different shard of the massive dataset, with gradients aggregated across workers to scale throughput. Option C is correct because at multi-GPU scale the input pipeline can easily starve the accelerators; using tf.data.Dataset with parallel reads (num_parallel_reads / interleave) and prefetching (prefetch(AUTOTUNE)) overlaps data loading and preprocessing with GPU computation so the GPUs are not idle. Option D is correct because synchronous gradient updates (e.g., all-reduce via NCCL, as in tf.distribute.MirroredStrategy) aggregate gradients from all workers before applying the optimizer step, which keeps replicas consistent and yields stable, reproducible convergence for deep learning training. Option A is not correct because a single large batch size is not itself an essential distributed-training strategy; batch size is a tuning choice, and naively enlarging it can hurt convergence and may not fit in memory. Option E is not correct because asynchronous gradient updates are an alternative to, not a requirement for, efficient distributed training, and they can introduce stale gradients that degrade model accuracy and reproducibility.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a single large batch size across all workers.
Why it's wrong here
One global batch size cannot be split across workers without partitioning; each worker needs its own per-device batch, with gradients synchronised each step. It is tempting because large batches do improve throughput on a single GPU, so it would be right when tuning one device's memory and convergence rather than scaling across many.
- ✓
Implement data parallelism.
Why this is correct
Data parallelism replicates the model across GPUs and splits each batch between them, satisfying the massive-dataset constraint by scaling throughput linearly with device count. Gradients are then aggregated, letting all GPUs train the same model concurrently.
- ✓
Ensure that the input pipeline is not a bottleneck by using tf.data.Dataset with prefetching and parallel reads.
Why this is correct
Feeding multiple GPUs requires the input pipeline to sustain their combined consumption; tf.data.Dataset with prefetching and parallel reads overlaps data loading with computation, preventing I/O stalls that would otherwise leave GPUs idle and negate the distributed training speed-up.
- ✓
Use synchronous gradient updates.
Why this is correct
Synchronous gradient updates aggregate all workers' gradients before each step, keeping replicas consistent and producing convergence behaviour equivalent to larger-batch training. This satisfies the multi-GPU constraint by avoiding the stale-gradient divergence that asynchronous updates introduce.
- ✗
Use asynchronous gradient updates to reduce communication overhead.
Why it's wrong here
Asynchronous updates let workers proceed on stale gradients, which harms convergence and reproducibility; synchronous all-reduce is what keeps replicas consistent. It is tempting because async does cut communication stalls, so it would be correct for latency-tolerant parameter-server setups, not for tightly coupled multi-GPU training needing stable accuracy.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.