mediumMultiple Choice
PMLE Practice Question: A data scientist trained a model on a single GPU…
A data scientist trained a model on a single GPU but needs to train on multiple GPUs for a larger dataset. They observe that training time does not decrease linearly with additional GPUs. Which common issue is most likely?
⚠ Common exam trap
PMLE often tests the assumption that adding more GPUs always linearly reduces training time; candidates must consider Amdahl's law and identify bottlenecks like data loading or communication overhead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Data pipeline bottleneck.
The most likely issue is a data pipeline bottleneck. When training on multiple GPUs, if the data input pipeline cannot feed data fast enough, the GPUs will idle waiting for data, leading to sublinear scaling. This is a common problem in distributed training where I/O or preprocessing becomes the limiting factor.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Overfitting.
Why it's wrong here
Overfitting concerns generalisation, not scaling efficiency; it does not explain sublinear speedup. The real cause is communication and synchronisation overhead between GPUs, such as gradient all-reduce. Overfitting is tempting because it is a familiar training pathology, but it would show as a validation-loss gap.
- ✗
Model architecture too simple.
Why it's wrong here
A simple architecture reduces compute per step but does not cause sublinear multi-GPU scaling; distributed communication overhead does. Simplicity is tempting when diagnosing underutilisation, yet the symptom here is scaling efficiency, not model capacity or convergence quality.
- ✗
Learning rate too high.
Why it's wrong here
A high learning rate affects convergence stability, not parallel scaling efficiency; sublinear speedup stems from inter-GPU communication and synchronisation overhead. Learning rate is tempting because it is a common training failure, but it would manifest as divergence or oscillation rather than diminishing returns per added GPU.
- ✓
Data pipeline bottleneck.
Why this is correct
Scaling GPUs only speeds up the compute graph; if the input pipeline cannot feed data fast enough, accelerators idle between steps. The bottleneck lies in data loading and preprocessing, so adding GPUs yields sublinear gains.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.