hardMultiple Choice
PMLE Practice Question: A data engineer is troubleshooting a Vertex AI…
A data engineer is troubleshooting a Vertex AI Endpoint that serves a large BERT model. After deployment, many prediction requests fail with 'Out of Memory' errors. The machine type is n1-standard-8 (30 GB memory) with no accelerator. Which action will most likely resolve the issue?
⚠ Common exam trap
PMLE often tests the misconception that adding a GPU fixes OOM errors, when OOM is a host memory issue and the correct fix is a high-memory machine type.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Change the machine type to n1-highmem-16 (104 GB memory).
The Out of Memory errors occur because the n1-standard-8 has only 30 GB RAM, which is insufficient to load the large BERT model plus runtime overhead. Switching to n1-highmem-16 provides 104 GB RAM, giving ample headroom for the model and inference batch, directly resolving the OOM condition.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Change the machine type to n1-highmem-16 (104 GB memory).
Why this is correct
The n1-standard-8's 30 GB cannot hold the large BERT model plus inference tensors, causing out-of-memory failures. Moving to n1-highmem-16 raises memory to 104 GB, satisfying the model's footprint without accelerators. More vCPUs alone would not increase memory proportionally.
- ✗
Use batch prediction instead of online prediction.
Why it's wrong here
Batch prediction processes inputs asynchronously from Cloud Storage and returns results to a bucket, so it cannot serve the live endpoint requests failing with out-of-memory errors. It is correct for large offline scoring jobs, not for online serving that requires immediate responses.
- ✗
Add a GPU accelerator (e.g., NVIDIA T4) to offload computation.
Why it's wrong here
GPU helps compute but does not increase available RAM.
- ✗
Quantize the model from FP32 to INT8.
Why it's wrong here
Quantisation shrinks weights and activations, cutting memory use, but it changes numerical precision and requires re-exporting or re-converting the model, which the stem does not mention as available. It is tempting because INT8 is the standard remedy for memory-bound inference, and would be correct if the model could be re-quantised and accuracy loss were acceptable.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.