mediumMultiple Choice
PMLE Practice Question: Training a large neural network on Vertex AI and…
A company is training a large neural network on Vertex AI and training jobs keep failing with 'Out of memory' errors. The VM uses a standard n1-standard-4 machine with 15 GB RAM. Which action should they take first?
⚠ Common exam trap
The trap here is that candidates often jump to scaling up infrastructure (larger machine or distributed training) instead of first tuning the training hyperparameter (batch size) that directly controls memory consumption, which is the simplest and most cost-effective fix.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reduce the batch size in the training script
The 'Out of memory' error on a n1-standard-4 VM (15 GB RAM) indicates the model's memory footprint exceeds available RAM. Reducing the batch size directly decreases the memory required for storing intermediate activations and gradients during training, which is the most immediate and cost-effective fix without changing the underlying infrastructure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger machine type like n1-standard-16
Why it's wrong here
A larger machine type adds vCPUs and RAM, but the out-of-memory error stems from GPU memory, which n1-standard-16 does not increase. It is tempting because more RAM helps CPU-bound workloads, and would be correct when host memory, not accelerator memory, is exhausted.
- ✓
Reduce the batch size in the training script
Why this is correct
Out-of-memory errors arise when the training batch exceeds available GPU or host RAM. Reducing batch size lowers peak memory per step, allowing the model to fit within the n1-standard-4's 15 GB. This is the fastest, least disruptive first remedy before considering larger machines.
- ✗
Enable distributed training across multiple VMs
Why it's wrong here
Distributed training splits work across VMs but each replica still holds the model and activations, so per-machine memory exhaustion persists. It is tempting because it scales throughput for large datasets, and would be correct when training time, not single-VM memory, is the bottleneck.
- ✗
Switch the training to CPU only
Why it's wrong here
CPU-only training still loads the same model and activations into the same 15 GB RAM, so the out-of-memory error recurs, just far slower. It is tempting to avoid GPU memory limits, and would be correct for small models or workloads where GPU acceleration is unavailable.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.