PMLE Scaling Prototypes into ML Models Practice Question
An ML engineer is preparing to train a large recommendation model on Vertex AI. The model uses a custom training loop in PyTorch and requires a multi-node cluster with 8 A100 GPUs per node. The engineer wants to minimize training time and ensure the job can recover from a node failure without restarting from scratch. Which combination of Vertex AI features should the engineer use?
⚠ Common exam trap
Watch out — candidates often confuse orchestration or tuning features with the distributed training runtime, when only a properly configured custom training job with multiple replicas provides the required multi-node GPU cluster.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a custom training job with a worker pool of 8 A100 GPUs per replica and multiple replicas, enable distributed training with the PyTorch distributed launcher, and write checkpoints to Cloud Storage periodically.
The requirement is a multi-node A100 cluster with fault tolerance. A Vertex AI custom training job with multiple replicas, each with 8 A100 GPUs, supplies the cluster, and the PyTorch distributed launcher coordinates training across nodes. Writing checkpoints to Cloud Storage at intervals lets the job resume after a node failure without losing all progress, which is the standard pattern for large distributed training on Vertex AI.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a custom training job with a worker pool of 8 A100 GPUs per replica and multiple replicas, enable distributed training with the PyTorch distributed launcher, and write checkpoints to Cloud Storage periodically.
Why this is correct
A custom training job with multiple replicas and 8 A100 GPUs each provides the required multi-node cluster. Vertex AI handles the distributed environment variables and network setup, and the PyTorch distributed launcher coordinates ranks across nodes. Periodic checkpoints to Cloud Storage allow the job to resume from the latest checkpoint after a node failure, satisfying both performance and recovery requirements.
- ✗
Use Vertex AI Pipelines to orchestrate the training as a series of steps, with each step training on a subset of data, and use pipeline caching to avoid retraining on failure.
Why it's wrong here
Vertex AI Pipelines is an orchestration layer, not a distributed training runtime. Splitting training into steps on data subsets changes the training semantics and does not provide multi-node GPU communication. Pipeline caching avoids re-executing successful steps but cannot resume a partially completed distributed training run. This option does not meet the requirement for a multi-node A100 cluster or mid-run fault tolerance.
- ✗
Use a Vertex AI custom training job with a single replica of 8 A100 GPUs and enable automatic machine restart, then rely on the default checkpointing behavior of the PyTorch training loop.
Why it's wrong here
A single replica of 8 GPUs does not provide the required multi-node cluster, and automatic machine restart restarts the job from the beginning rather than resuming from a checkpoint. PyTorch training loops do not checkpoint by default unless the code explicitly saves state. This option fails both the cluster-size requirement and the recovery requirement, so it cannot deliver the desired outcome.
- ✗
Use a Vertex AI hyperparameter tuning job with a single replica of 8 A100 GPUs, relying on early stopping to reduce training time.
Why it's wrong here
Hyperparameter tuning runs many independent trials and does not provide multi-node distributed training within a single trial. A single replica of 8 GPUs cannot meet the multi-node requirement. Early stopping reduces wasted trials but does not enable recovery from a node failure mid-trial. This option addresses a different problem and does not satisfy the cluster size or fault-tolerance needs.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.