Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

Your team is training a very large transformer model that does not fit on a single GPU. They are using Vertex AI custom training with PyTorch. Which distributed training approach should they use?

⚠ Common exam trap

Google Cloud often tests the distinction between data parallelism (which replicates the model) and model parallelism (which splits the model), and the trap here is that candidates assume any distributed training framework (like DDP or Horovod) can handle oversized models, ignoring the fundamental memory constraint that data parallelism cannot overcome.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Model parallelism using pipeline parallelism

When a transformer model is too large to fit on a single GPU, model parallelism (specifically pipeline parallelism) is required because it splits the model's layers across multiple devices, with each device holding a subset of the model's parameters. Data parallelism (DDP) replicates the entire model on each GPU, which fails if the model exceeds a single GPU's memory. Pipeline parallelism allows training very large models by partitioning the model into stages and passing activations and gradients sequentially between devices.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Data parallelism using PyTorch DistributedDataParallel (DDP)

    Why it's wrong here

    DDP replicates the model on each GPU; model must fit on one GPU.

  • ✗

    Horovod with allreduce

    Why it's wrong here

    Horovod's allreduce ring replicates the full model on every worker, so a model exceeding one GPU's memory still cannot be held. It suits scaling training across many GPUs when each replica fits in device memory, not sharding an oversized model.

  • ✓

    Model parallelism using pipeline parallelism

    Why this is correct

    Pipeline parallelism splits the transformer's layers across GPUs, so each device holds only a subset of parameters and activations. This addresses the constraint that the model exceeds single-GPU memory, unlike data parallelism, which replicates the full model on every device.

  • ✗

    Vertex AI distributed training with TF_CONFIG

    Why it's wrong here

    TF_CONFIG is the TensorFlow distributed strategy configuration mechanism; it does not drive PyTorch training and offers no parameter sharding. It would be correct for TensorFlow Estimator or Keras distributed jobs on Vertex AI, not this PyTorch transformer.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

8 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You have a very large language model that does not fit on a single GPU. You need to train it efficiently across multiple GPUs on a single machine. Which approach should you use?

hard
  • A.Data parallelism with MirroredStrategy
  • B.Data parallelism with MultiWorkerMirroredStrategy
  • C.Use TPU training as TPUs have more memory
  • ✓ D.Model parallelism using pipeline parallelism

Why D: When a model does not fit on a single GPU, model parallelism is required because it partitions the model itself across devices. Pipeline parallelism is a specific model parallelism technique that splits the model into stages across GPUs and pipelines micro-batches to maintain utilization, making it the appropriate approach for training a very large model on multiple GPUs in one machine.

Variation 2. A team is training a large TensorFlow model that requires more memory than a single GPU provides. They have access to multiple GPUs on a single machine. Which distributed training strategy should they use to split the model layers across GPUs?

medium
  • A.tf.distribute.experimental.MultiWorkerMirroredStrategy
  • B.tf.distribute.experimental.ParameterServerStrategy
  • ✓ C.Manual device placement using tf.device to assign layers to specific GPUs
  • D.tf.distribute.MirroredStrategy

Why C: When a single model's layers exceed one GPU's memory, the model itself must be partitioned across devices — this is model parallelism. Manual device placement with tf.device('/GPU:0'), tf.device('/GPU:1'), etc. is the TensorFlow-native way to assign specific layers or operations to specific GPUs, splitting the model across them.

Variation 3. You are designing a distributed training job for a very large neural network that does not fit on a single machine. You need to split the model across multiple devices. Which TWO techniques can you use?

medium
  • A.ParameterServerStrategy
  • ✓ B.Pipeline parallelism
  • ✓ C.Operator-level model parallelism
  • D.Data parallelism with MirroredStrategy
  • E.MultiWorkerMirroredStrategy

Why B: Pipeline parallelism (B) is correct because it splits the model's layers into sequential stages placed on different devices, so a model too large for one machine can be distributed across devices while micro-batches flow through the pipeline. Operator-level model parallelism (C) is also correct because it partitions individual operations (e.g., splitting a large matrix multiplication or a layer's weights) across devices, which directly addresses a model that does not fit on a single machine. ParameterServerStrategy (A) is a data-parallel approach that replicates the full model on each worker and only shards the parameter updates, so it does not solve the memory problem of a model too large for one device. Data parallelism with MirroredStrategy (D) likewise replicates the entire model on every device, which is infeasible when the model does not fit on one machine. MultiWorkerMirroredStrategy (E) is also a data-parallel strategy that requires the full model to fit on each worker, so it does not satisfy the requirement.

Variation 4. A team is fine-tuning a large language model (LLaMA 2) using Vertex AI with a custom container on a multi-node GPU cluster. They need to implement model parallelism to fit the model across multiple GPUs because it does not fit into a single GPU memory. Which distributed training strategy should they use?

hard
  • A.Use Vertex AI Hyperparameter Tuning to find optimal model partitioning
  • B.Use tf.distribute.MirroredStrategy across all GPUs
  • ✓ C.Implement pipeline parallelism by manually splitting the model layers across GPUs and using a framework like PyTorch's RPC or Megatron-LM
  • D.Use Vertex AI distributed training with TF_CONFIG to set up multi-worker mirrored strategy and rely on XLA to partition the model

Why C: When a model does not fit into a single GPU's memory, model parallelism is required. Pipeline parallelism splits the model layers across multiple GPUs, and frameworks like PyTorch's RPC or Megatron-LM provide the necessary primitives to implement it. This is the correct strategy for fitting a large model like LLaMA 2 across multiple GPUs.

Variation 5. You are fine-tuning a large language model (LLM) from Hugging Face Transformers using Vertex AI Training. The model has 7 billion parameters and does not fit into the memory of a single GPU. You need to train across multiple GPUs, splitting the model layers across devices. Which distributed training approach should you use?

hard
  • ✓ A.Model parallelism using pipeline parallelism
  • B.Data parallelism with MultiWorkerMirroredStrategy
  • C.Mixed precision training (FP16)
  • D.Data parallelism with tf.distribute.MirroredStrategy

Why A: When a model is too large to fit on a single GPU, model parallelism is required to split the model's layers across multiple devices. Pipeline parallelism is a specific form of model parallelism that partitions layers into stages and pipelines micro-batches across devices, enabling training of models like a 7B-parameter LLM across multiple GPUs. This is the correct approach when memory, not throughput, is the binding constraint.

Variation 6. A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?

hard
  • A.Using Vertex AI Vizier to optimize the model parallelism strategy
  • B.Enabling Vertex AI AutoML to automatically distribute the model
  • ✓ C.Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe
  • ✓ D.A custom container with the distributed framework (e.g., PyTorch DDP) installed
  • ✓ E.Setting the --worker-machine-count flag when submitting the job

Why C: Pipeline parallelism requires explicit implementation in the training script, such as using `torch.distributed.pipeline.sync.Pipe` in PyTorch, to split the model layers across multiple GPUs. This is necessary when the model does not fit on a single GPU, and Vertex AI does not automatically handle model parallelism—it must be coded by the user.

Variation 7. A machine learning team is training a large transformer model on Vertex AI. They need to reduce training time by utilizing multiple GPUs across nodes, but the model is too large to fit into a single GPU memory. Which distributed training strategy should they use?

medium
  • ✓ A.Model parallelism using tf.distribute.experimental.PipelineMirroredStrategy
  • B.Data parallelism using tf.distribute.MirroredStrategy
  • C.Multi-worker mirrored strategy with a single worker per node
  • D.Hyperparameter tuning with Vertex AI Vizier

Why A: PipelineMirroredStrategy combines model parallelism (splitting the transformer layers across multiple GPUs) with pipeline parallelism to handle models that exceed single GPU memory. This strategy partitions the model into stages, each placed on a different GPU, and uses micro-batching to keep all GPUs busy, which is essential for large transformer models that cannot fit into a single GPU's memory.

Variation 8. A data scientist is fine-tuning a large language model from Hugging Face using Vertex AI Training with a GPU. The model has 7 billion parameters and does not fit on a single GPU. They need to split the model across multiple GPUs and train with data parallelism. Which strategy should they use?

hard
  • A.Use Vertex AI's AutoML to automatically distribute the model.
  • ✓ B.Use pipeline parallelism via a custom container with DeepSpeed and data parallelism across workers using PyTorch DDP, configured with Vertex AI distributed training.
  • C.Use Vertex AI's hyperparameter tuning with multiple trials.
  • D.Configure a multi-worker mirrored strategy with TensorFlow, setting TF_CONFIG to use all GPUs on each worker.

Why B: It combines pipeline parallelism (via DeepSpeed) to split the 7B-parameter model across multiple GPUs, with data parallelism (via PyTorch DDP) to replicate the model across workers for training on larger batches. Vertex AI distributed training coordinates the multi-worker setup, making this the only viable strategy for a model that exceeds single-GPU memory while requiring data parallelism.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.