mediumMultiple Choice
PMLE Practice Question: A team has a prototype image classification model…
A team has a prototype image classification model trained on a small dataset using TensorFlow Keras on a single GPU. They need to train on a larger dataset (1 million images) using a distributed strategy on Vertex AI with 8 GPUs. They implement a MirroredStrategy for data parallelism. During the first few epochs, the training speed does not improve significantly compared to a single GPU, and GPU utilization is low. The data is stored as JPEG files in Cloud Storage, and the input pipeline uses tf.data with map to decode images. What is the most likely cause?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The data loading from Cloud Storage is a bottleneck.
Reading and decoding JPEG images from Cloud Storage can be I/O-bound, causing low GPU utilization. Even though MirroredStrategy is used, if the input pipeline cannot supply data fast enough, GPUs will spend time waiting for data. Option A is incorrect because a large batch size per GPU typically increases memory usage and computational load, not decreases utilization; it might even improve utilization if the model fits. Option B is incorrect because MirroredStrategy is a straightforward configuration for synchronous data parallelism and does not require complex tuning; it is likely configured correctly. Option D is incorrect because the model size does not directly cause low GPU utilization; even a small model can saturate GPUs if the data pipeline is efficient. The primary bottleneck here is the data loading from Cloud Storage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The batch size per GPU is too large.
Why it's wrong here
Larger per-GPU batches would raise utilisation, not depress it, and MirroredStrategy scales the global batch across replicas anyway. Batch size tuning is the right lever when GPUs are saturated but convergence stalls; here the bottleneck is host-side JPEG decoding inside map, which starves every replica.
- ✗
The MirroredStrategy is not properly configured.
Why it's wrong here
MirroredStrategy replicates variables across all eight GPUs and is the standard single-machine data-parallel strategy, so misconfiguration would raise errors rather than silently halve throughput. It is tempting because low GPU utilisation suggests a distribution fault, and would be correct if devices were undiscovered or the strategy were applied outside the strategy scope.
- ✓
The data loading from Cloud Storage is a bottleneck.
Why this is correct
Decoding JPEGs inside tf.data.map runs on CPU and reads from Cloud Storage over the network, starving eight GPUs. MirroredStrategy replicates compute but not input throughput, so low GPU utilisation persists until parallel interleave and prefetch are added.
- ✗
The model is too small for distributed training.
Why it's wrong here
Model size does not gate MirroredStrategy, which replicates any graph across the eight GPUs. Small models are the correct concern when per-step communication overhead dominates compute, so scaling out adds latency; here the bottleneck is the undecoded JPEG input pipeline, not parameter count.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.