Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You have a very large language model that does not fit on a single GPU. You need to train it efficiently across multiple GPUs on a single machine. Which approach should you use?

⚠ Common exam trap

PMLE often tests the misconception that data parallelism can handle models too large for one GPU, when in fact data parallelism replicates the model and only model parallelism (including pipeline parallelism) partitions it across devices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Model parallelism using pipeline parallelism

When a model does not fit on a single GPU, model parallelism is required because it partitions the model itself across devices. Pipeline parallelism is a specific model parallelism technique that splits the model into stages across GPUs and pipelines micro-batches to maintain utilization, making it the appropriate approach for training a very large model on multiple GPUs in one machine.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Data parallelism with MirroredStrategy

    Why it's wrong here

    MirroredStrategy replicates the full model on each GPU, so a model exceeding single-GPU memory still cannot fit; it only splits the input batch. It is tempting because data parallelism scales throughput well when the model fits per device, which is not this scenario.

  • ✗

    Data parallelism with MultiWorkerMirroredStrategy

    Why it's wrong here

    Data parallelism replicates the full model on every GPU, so a model too large for one device still will not fit. It is tempting because MultiWorkerMirroredStrategy scales training across workers, and would be correct if the model fitted per GPU and throughput, not memory, were the constraint.

  • ✗

    Use TPU training as TPUs have more memory

    Why it's wrong here

    Switching to TPUs requires rewriting the training stack for XLA and does not address the stated single-machine multi-GPU requirement; TPU memory is not a configuration setting. It tempts because TPUs offer large high-bandwidth memory for very large models in cloud environments.

  • ✓

    Model parallelism using pipeline parallelism

    Why this is correct

    Pipeline parallelism splits the model's layers across GPUs so each device holds only a subset of parameters, with activations passed between stages. This fits a model too large for one GPU while training efficiently within a single machine.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.