PMLE Scaling Prototypes into ML Models Practice Question
Your Vertex AI custom training job is failing with an out-of-memory error on a single GPU. You need to reduce memory usage without changing the model architecture. Which approach should you try first?
⚠ Common exam trap
PMLE often tests the order of operations for troubleshooting OOM, trapping candidates who jump to advanced techniques like mixed precision or model parallelism instead of the simplest, most direct fix: reducing batch size.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Decrease the batch size
Decreasing the batch size directly reduces the memory required for activations and gradients, making it the simplest and most immediate way to resolve out-of-memory errors without altering the model architecture. It is a standard first step because it requires no code changes beyond a hyperparameter and often resolves OOM with minimal impact on convergence if adjusted with learning rate.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Decrease the batch size
Why this is correct
Reducing the batch size lowers the number of samples held in GPU memory per step, directly cutting activation and gradient memory consumption. It requires no architecture change, satisfying the stem's constraint, and is the least invasive first remedy before considering gradient checkpointing or mixed precision.
- ✗
Implement model parallelism across GPUs
Why it's wrong here
Implementing model parallelism across GPUs is incorrect because the out-of-memory error occurs on a *single* GPU. This approach requires multiple GPUs to distribute the model's layers, which does not reduce the memory footprint on the *initial* failing GPU. It is tempting as it addresses OOM errors for models too large to fit into the memory of *any single GPU*, enabling training by spreading the model's components across several devices.
- ✗
Use gradient accumulation
Why it's wrong here
Gradient accumulation reduces peak activation memory only when the per-device batch is split into micro-batches, but it does not shrink the memory consumed by the model's parameters, optimiser states or a single forward/backward pass. It is genuinely useful for simulating larger effective batch sizes on limited hardware.
- ✗
Enable mixed precision training (FP16)
Why it's wrong here
Mixed precision training reduces memory by storing weights and activations in FP16, but the out-of-memory error on a single GPU stems from the model exceeding the GPU’s total VRAM capacity, not from precision overhead. FP16 can even increase memory pressure if the model requires FP32 master weights for gradient updates. It is tempting because mixed precision is commonly used to speed up training and reduce memory on modern GPUs with Tensor Cores, and would be correct if the goal were to accelerate computation or fit a slightly larger batch size on a memory-constrained GPU.
Visual reference
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.