Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?

⚠ Common exam trap

PMLE often tests the misconception that data parallelism solves memory constraints, when in fact it replicates the model and only model/pipeline parallelism addresses models too large for one GPU.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Model parallelism using pipeline parallelism

When a model is too large to fit on a single GPU, model parallelism is required to split the model's layers across multiple devices. Pipeline parallelism is a specific form of model parallelism that partitions layers into stages and pipelines micro-batches across devices, enabling training of models like a 7B-parameter LLM across multiple GPUs. This is the correct approach when memory, not throughput, is the binding constraint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Model parallelism using pipeline parallelism

    Why this is correct

    Pipeline parallelism splits the model's layers across GPUs, with each device holding a subset and passing activations onward. This addresses the constraint that seven billion parameters exceed single-GPU memory, unlike data parallelism which replicates the full model per device.

  • ✗

    Data parallelism with MultiWorkerMirroredStrategy

    Why it's wrong here

    MultiWorkerMirroredStrategy replicates the full model on every worker, so a 7B-parameter model still exceeds each GPU's memory; it splits data, not layers. It is the correct choice when the model fits per device and you need to scale batch throughput across replicas.

  • ✗

    Mixed precision training (FP16)

    Why it's wrong here

    FP16 halves activation and weight storage but does not partition the model, so layers still reside wholly on one GPU and the memory shortfall persists. Mixed precision is the right optimisation when the model already fits and you want faster training with reduced memory footprint.

  • ✗

    Data parallelism with tf.distribute.MirroredStrategy

    Why it's wrong here

    MirroredStrategy is single-machine, synchronous data parallelism that copies the entire model to each GPU, leaving the 7B parameters undivided and still too large. It suits multi-GPU training on one host when the model fits per device and throughput must scale.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.