PMLE Scaling Prototypes into ML Models Practice Question
A machine learning engineer is training a TensorFlow model on Vertex AI using distributed training with the MultiWorkerMirroredStrategy. The training job uses 4 workers with 4 GPUs each. The engineer notices that the training is not scaling linearly. What is the most likely cause?
⚠ Common exam trap
Test-takers frequently assume more workers always means linear speedup, ignoring the fixed overhead of gradient synchronization that becomes the dominant factor in distributed training.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Communication overhead due to gradient synchronization
With MultiWorkerMirroredStrategy, each worker computes gradients independently on its local batch, then all-reduces gradients across workers via collective communication (e.g., NCCL or gRPC). As the number of workers increases, the communication overhead for gradient synchronization grows, often dominating the per-step time and preventing linear scaling. This is the most common bottleneck in distributed TensorFlow training, especially with many workers or small batch sizes per worker.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model architecture is too simple to benefit from distribution
Why it's wrong here
A simple model still scales if communication overhead is low; architecture complexity does not determine MultiWorkerMirroredStrategy efficiency. This option would fit a scenario where the model has too few parameters to occupy the GPUs, not one where four workers with four GPUs each underperform.
- ✗
The workers are not using the same version of TensorFlow
Why it's wrong here
Version mismatches cause job startup failures or collective-operation errors, not gradual scaling loss. Uniform TensorFlow versions are a prerequisite for launching MultiWorkerMirroredStrategy at all, so this would be the answer when workers refuse to form a cluster.
- ✓
Communication overhead due to gradient synchronization
Why this is correct
MultiWorkerMirroredStrategy performs all-reduce gradient synchronisation across workers each step, so inter-worker communication overhead grows with worker count and prevents linear scaling. With 4 workers and 16 GPUs, network latency and bandwidth for exchanging gradients, not compute, become the bottleneck limiting throughput gains.
- ✗
The GPUs are not configured correctly
Why it's wrong here
Misconfigured GPUs would typically cause job failure or errors, not merely sublinear scaling. GPU configuration matters when devices are absent from the visible device list or TensorFlow cannot allocate them, which is a setup fault rather than a scaling bottleneck.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.