PMLE Scaling Prototypes into ML Models Practice Question
A team is training a large TensorFlow model that requires more memory than a single GPU provides. They have access to multiple GPUs on a single machine. Which distributed training strategy should they use to split the model layers across GPUs?
⚠ Common exam trap
The trap is assuming any 'distributed strategy' solves memory limits — most strategies (Mirrored, MultiWorker, ParameterServer) replicate the model and only help with speed, not with fitting an oversized model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Manual device placement using tf.device to assign layers to specific GPUs
When a single model's layers exceed one GPU's memory, the model itself must be partitioned across devices — this is model parallelism. Manual device placement with tf.device('/GPU:0'), tf.device('/GPU:1'), etc. is the TensorFlow-native way to assign specific layers or operations to specific GPUs, splitting the model across them.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
tf.distribute.experimental.MultiWorkerMirroredStrategy
Why it's wrong here
MultiWorkerMirroredStrategy replicates the full model on each worker and synchronises gradients, so it cannot split layers when the model exceeds one GPU's memory. It is tempting because it scales across multiple workers, and it would be correct for data-parallel training across several machines.
- ✗
tf.distribute.experimental.ParameterServerStrategy
Why it's wrong here
ParameterServerStrategy places variables on parameter servers and is designed for multi-machine clusters, not splitting layers across GPUs within one machine. It is tempting because it distributes training, and it would be correct when workers and parameter servers span several hosts on a network.
- ✓
Manual device placement using tf.device to assign layers to specific GPUs
Why this is correct
tf.device lets you pin individual layers to named GPUs, so a model too large for one device is partitioned layer-by-layer across the machine's GPUs. This directly satisfies the stem's requirement to split model layers rather than replicate the model.
- ✗
tf.distribute.MirroredStrategy
Why it's wrong here
MirroredStrategy replicates identical variables on every GPU and synchronises updates, so each device must hold the whole model; it cannot partition layers. It is tempting because it uses all GPUs on one machine, and it would be correct when the model fits within a single GPU's memory.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.