PMLE Scaling Prototypes into ML Models Practice Question
A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?
⚠ Common exam trap
PMLE often tests the confusion between data parallelism (replicating the model) and model parallelism (splitting the model), causing candidates to choose MirroredStrategy when the model does not fit on a single GPU.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM
When a model does not fit into a single GPU's memory, model parallelism is required. Pipeline parallelism splits the model layers across multiple GPUs, and frameworks like PyTorch's RPC or Megatron-LM provide the necessary primitives to implement it. This is the correct strategy for fitting a large model like LLaMA 2 across multiple GPUs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Vertex AI Hyperparameter Tuning to find optimal model partitioning
Why it's wrong here
Hyperparameter Tuning searches training configurations such as learning rate or batch size; it does not partition model layers across GPUs. It is correct when optimising training metrics within a fixed architecture. The stem requires distributing model parameters across devices, which tuning cannot accomplish.
- ✗
Use tf.distribute.MirroredStrategy across all GPUs
Why it's wrong here
MirroredStrategy replicates the full model on every GPU and synchronises gradients, so it cannot fit a model exceeding single-GPU memory. It suits data-parallel workloads where the model fits per device. The stem requires sharding model layers across devices, which mirrored replication does not perform.
- ✓
Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM
Why this is correct
Pipeline parallelism partitions consecutive model layers into stages across GPUs, so each device holds only a fraction of parameters. This directly addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.
- ✗
Use Vertex AI distributed training with TF_CONFIG to set up multi-worker mirrored strategy and rely on XLA to partition the model
Why it's wrong here
Multi-worker mirrored strategy still replicates the complete model on each worker, and XLA does not automatically partition a model across devices for memory relief. This suits data-parallel training with per-replica fitting models. The stem needs explicit parameter or pipeline sharding, which mirrored replication cannot provide.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.