PMLE Scaling Prototypes into ML Models Practice Question
A research team is training a very large Transformer model that does not fit into the memory of a single GPU. They have access to multiple GPUs on a single machine and want to split the model layers across GPUs. Which distributed training strategy should they use?
⚠ Common exam trap
Many candidates confuse data parallelism (which replicates the model) with model parallelism (which splits the model), and overlooking that the problem explicitly states the model does not fit in a single GPU's memory, making any data-parallel strategy incorrect.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pipeline parallelism (model parallelism)
Pipeline parallelism (model parallelism) is the correct choice because it splits the model's layers across multiple GPUs, allowing a model too large for a single GPU's memory to be trained. Each GPU holds a subset of layers and processes micro-batches in a pipelined fashion, enabling training of models that exceed single-device memory. This directly addresses the memory constraint by distributing the model parameters, not just the data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
MultiWorkerMirroredStrategy
Why it's wrong here
MultiWorkerMirroredStrategy replicates the full model on every worker and synchronises gradients, so each GPU must still hold all parameters, which the oversized Transformer cannot. It is tempting because it scales across multiple machines, and would be correct when the model fits per device but throughput demands more workers.
- ✗
Parameter server strategy
Why it's wrong here
Parameter server strategy distributes variables across servers and workers, but each worker still computes full forward and backward passes over complete layers, so it does not split the model itself. It is tempting because it scales to many nodes, and would be correct for large embedding tables or sparse recommendation models.
- ✗
MirroredStrategy (data parallelism)
Why it's wrong here
MirroredStrategy keeps an identical full replica of the model on each GPU and only shards the input batch, so per-GPU memory is unchanged and the model still cannot fit. It is tempting because it is the simplest single-machine multi-GPU path, and would be correct when the model fits on one device but training needs more throughput.
- ✓
Pipeline parallelism (model parallelism)
Why this is correct
Pipeline parallelism partitions the model's layers across GPUs, so each device holds only a subset of parameters and activations. This directly addresses the constraint that the model exceeds single-GPU memory. Data parallelism would replicate the full model on every GPU, which remains impossible.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.