PMLE Scaling Prototypes into ML Models Practice Question
You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?
⚠ Common exam trap
PMLE often tests the misconception that data parallelism solves memory constraints, when in fact it replicates the model and only model/pipeline parallelism addresses models too large for one GPU.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model parallelism using pipeline parallelism
When a model is too large to fit on a single GPU, model parallelism is required to split the model's layers across multiple devices. Pipeline parallelism is a specific form of model parallelism that partitions layers into stages and pipelines micro-batches across devices, enabling training of models like a 7B-parameter LLM across multiple GPUs. This is the correct approach when memory, not throughput, is the binding constraint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Model parallelism using pipeline parallelism
Why this is correct
Pipeline parallelism splits the model's layers across GPUs, with each device holding a subset and passing activations onward. This addresses the constraint that seven billion parameters exceed single-GPU memory, unlike data parallelism which replicates the full model per device.
- ✗
Data parallelism with MultiWorkerMirroredStrategy
Why it's wrong here
MultiWorkerMirroredStrategy replicates the full model on every worker, so a 7B-parameter model still exceeds each GPU's memory; it splits data, not layers. It is the correct choice when the model fits per device and you need to scale batch throughput across replicas.
- ✗
Mixed precision training (FP16)
Why it's wrong here
FP16 halves activation and weight storage but does not partition the model, so layers still reside wholly on one GPU and the memory shortfall persists. Mixed precision is the right optimisation when the model already fits and you want faster training with reduced memory footprint.
- ✗
Data parallelism with tf.distribute.MirroredStrategy
Why it's wrong here
MirroredStrategy is single-machine, synchronous data parallelism that copies the entire model to each GPU, leaving the 7B parameters undivided and still too large. It suits multi-GPU training on one host when the model fits per device and throughput must scale.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.