hardMultiple Choice
PDE Practice Question: A team is training a large model using a custom…
A team is training a large model using a custom container with TensorFlow on Vertex AI Training. They need to use multiple GPUs across several machines. Which strategy should they implement to maximize training throughput?
⚠ Common exam trap
Google often tests the distinction between single-machine multi-GPU strategies (MirroredStrategy) and multi-machine distributed strategies (MultiWorkerMirroredStrategy), leading candidates to pick D when they overlook the requirement for multiple machines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Vertex AI Training with a custom job specifying workerPoolSpecs and MultiWorkerMirroredStrategy
Vertex AI Training's custom job with workerPoolSpecs enables multi-machine, multi-GPU distributed training, and TensorFlow's MultiWorkerMirroredStrategy is specifically designed for synchronous distributed training across multiple workers. This combination maximizes throughput by efficiently synchronizing gradients across all GPUs on all machines using all-reduce communication, which is essential for large model training.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud TPU Pods for distributed training
Why it's wrong here
TPU Pods run XLA-compiled TensorFlow on TPU hardware, not NVIDIA GPUs, so they cannot host the team's GPU-based custom container across machines. Tempting because TPU Pods deliver the highest throughput for very large models, and would be correct if the container were TPU-compatible.
- ✗
Use Dataflow for distributed training
Why it's wrong here
Dataflow is a Apache Beam batch and streaming pipeline service with no GPU accelerators or distributed training runtime, so it cannot execute the TensorFlow training job. It tempts because it distributes data processing across many workers, and would be correct for preprocessing training data rather than training the model itself.
- ✓
Use Vertex AI Training with a custom job specifying workerPoolSpecs and MultiWorkerMirroredStrategy
Why this is correct
MultiWorkerMirroredStrategy performs all-reduce synchronisation across GPUs on every machine, giving true data-parallel training over the whole cluster. Specifying workerPoolSpecs in the custom job provisions those multiple machines, satisfying the stem's requirement for multi-machine, multi-GPU throughput rather than single-node scaling.
- ✗
Use a single worker with multiple GPUs and TensorFlow MirroredStrategy
Why it's wrong here
MirroredStrategy replicates variables across GPUs inside one worker, so it cannot span the several machines the stem requires. It tempts because it is the standard single-host multi-GPU approach, and would be correct if all GPUs sat in one machine rather than across a cluster.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.