Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?

⚠ Common exam trap

PMLE often tests the confusion between data parallelism (replicating the model) and model parallelism (splitting the model), causing candidates to choose MirroredStrategy when the model does not fit on a single GPU.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM

When a model does not fit into a single GPU's memory, model parallelism is required. Pipeline parallelism splits the model layers across multiple GPUs, and frameworks like PyTorch's RPC or Megatron-LM provide the necessary primitives to implement it. This is the correct strategy for fitting a large model like LLaMA 2 across multiple GPUs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Vertex AI Hyperparameter Tuning to find optimal model partitioning

    Why it's wrong here

    Hyperparameter Tuning searches training configurations such as learning rate or batch size; it does not partition model layers across GPUs. It is correct when optimising training metrics within a fixed architecture. The stem requires distributing model parameters across devices, which tuning cannot accomplish.

  • ✗

    Use tf.distribute.MirroredStrategy across all GPUs

    Why it's wrong here

    MirroredStrategy replicates the full model on every GPU and synchronises gradients, so it cannot fit a model exceeding single-GPU memory. It suits data-parallel workloads where the model fits per device. The stem requires sharding model layers across devices, which mirrored replication does not perform.

  • ✓

    Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM

    Why this is correct

    Pipeline parallelism partitions consecutive model layers into stages across GPUs, so each device holds only a fraction of parameters. This directly addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

  • ✗

    Use Vertex AI distributed training with TF_CONFIG to set up multi-worker mirrored strategy and rely on XLA to partition the model

    Why it's wrong here

    Multi-worker mirrored strategy still replicates the complete model on each worker, and XLA does not automatically partition a model across devices for memory relief. This suits data-parallel training with per-replica fitting models. The stem needs explicit parameter or pipeline sharding, which mirrored replication cannot provide.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.