Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are designing a distributed training job for a very large neural network that does not fit on a single machine. You need to split the model across multiple devices. Which TWO techniques can you use?

⚠ Common exam trap

The trap is confusing data parallelism with model parallelism; candidates may pick ParameterServerStrategy or MirroredStrategy because they are familiar distributed training strategies, but the exam expects recognition that only pipeline and operator-level parallelism split the model across devices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pipeline parallelism

Pipeline parallelism (B) is correct because it splits the model's layers into sequential stages placed on different devices, so a model too large for one machine can be distributed across devices while micro-batches flow through the pipeline. Operator-level model parallelism (C) is also correct because it partitions individual operations (e.g., splitting a large matrix multiplication or a layer's weights) across devices, which directly addresses a model that does not fit on a single machine. ParameterServerStrategy (A) is a data-parallel approach that replicates the full model on each worker and only shards the parameter updates, so it does not solve the memory problem of a model too large for one device. Data parallelism with MirroredStrategy (D) likewise replicates the entire model on every device, which is infeasible when the model does not fit on one machine. MultiWorkerMirroredStrategy (E) is also a data-parallel strategy that requires the full model to fit on each worker, so it does not satisfy the requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    ParameterServerStrategy

    Why it's wrong here

    ParameterServerStrategy distributes variables across parameter servers but each worker still holds a full model replica, so it does not partition the network itself. It is tempting because it scales to many workers, and it would be correct for large data-parallel training where the model fits on each worker.

  • ✓

    Pipeline parallelism

    Why this is correct

    Pipeline parallelism divides the model into sequential stages placed on different devices, with micro-batches flowing between them. This splits the model across devices, satisfying the stem's constraint that the neural network does not fit on a single machine.

  • ✓

    Operator-level model parallelism

    Why this is correct

    Operator-level model parallelism partitions individual operators, such as matrix multiplications, across devices so each holds a slice of the computation. This splits the model itself across devices, satisfying the stem's constraint that the network cannot fit on a single machine.

  • ✗

    Data parallelism with MirroredStrategy

    Why it's wrong here

    MirroredStrategy replicates the full model on each device and splits only the input data, so every replica must still hold the entire network. It is tempting because it scales throughput well, and it would be the right choice when the model fits on one device but training data volume demands parallel batches.

  • ✗

    MultiWorkerMirroredStrategy

    Why it's wrong here

    MultiWorkerMirroredStrategy also replicates complete model copies across workers, sharding the training data rather than the model itself. It is tempting because it scales across many hosts, and it would be correct when the model fits in device memory but training must span multiple machines for speed.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.