PMLE Scaling Prototypes into ML Models Practice Question
A data scientist is training a very large neural network using Vertex AI with multiple GPUs across multiple nodes. The model does not fit on a single GPU, so they need to use both data parallelism and model parallelism (pipeline parallelism). Which THREE components or configurations are required to set up distributed training with Vertex AI?
⚠ Common exam trap
This question tests the misconception that Vertex AI automatically handles model parallelism (e.g., via AutoML or Vizier), when in reality the user must manually implement it in the training script using frameworks like PyTorch or TensorFlow.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe
Pipeline parallelism requires explicit implementation in the training script, such as using `torch.distributed.pipeline.sync.Pipe` in PyTorch, to split the model layers across multiple GPUs. This is necessary when the model does not fit on a single GPU, and Vertex AI does not automatically handle model parallelism—it must be coded by the user.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Using Vertex AI Vizier to optimize the model parallelism strategy
Why it's wrong here
Vizier is for hyperparameter tuning, not for configuring parallelism.
- ✗
Enabling Vertex AI AutoML to automatically distribute the model
Why it's wrong here
AutoML is for automated model building, not for custom distributed training.
- ✓
Implementing pipeline parallelism manually in the training script using torch.distributed.pipeline.sync.Pipe
Why this is correct
Manual implementation of pipeline parallelism is required as Vertex AI does not provide built-in model parallelism.
- ✓
A custom container with the distributed framework (e.g., PyTorch DDP) installed
Why this is correct
Custom container needed to include the framework code and dependencies.
- ✓
Setting the --worker-machine-count flag when submitting the job
Why this is correct
Specifies the number of worker nodes for distributed training.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.