hardMultiple Choice
PMLE Practice Question: A company serves a PyTorch model using a custom…
A company serves a PyTorch model using a custom container on Vertex AI Prediction. They notice that after a few hours, the endpoint returns 502 errors. The logs show 'Out of memory' errors. The container has a memory limit of 4GB, and the model loads a 3GB vocabulary file. What is the most likely cause and best fix?
⚠ Common exam trap
Google Cloud often tests the misconception that OOM errors are always solved by increasing memory, but the trap here is that the real issue is inefficient resource reuse—loading a large file per request—rather than insufficient total memory.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Load the vocabulary file once at startup and reuse it.
The 502 errors and 'Out of memory' errors indicate that the container is running out of memory during inference. Since the model loads a 3GB vocabulary file, and the container has only 4GB of memory, loading this file repeatedly for each prediction request (e.g., inside the prediction handler) would quickly exhaust memory. The correct fix is to load the vocabulary file once at container startup and reuse it across all requests, which is a standard best practice for serving models with large static assets.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the container memory to 8GB.
Why it's wrong here
Raising the limit to 8GB leaves the 3GB vocabulary plus PyTorch runtime, activations and per-request tensors still competing within the same container, so OOM recurs under load. It is tempting because memory limits are the usual lever for OOM, and would be right if the model itself simply exceeded 4GB at load time.
- ✓
Load the vocabulary file once at startup and reuse it.
Why this is correct
Reloading the 3GB vocabulary per request exhausts the 4GB container limit, causing out-of-memory 502s. Loading it once at startup keeps it resident in memory and reused across requests, eliminating the repeated allocation that triggers the errors.
- ✗
Increase the number of replicas to distribute load.
Why it's wrong here
Adding replicas multiplies the number of containers each loading the 3GB vocabulary, so aggregate memory pressure rises and per-container OOM persists. The fix is reducing per-container memory footprint or raising the memory limit; horizontal scaling addresses throughput, not a single container exceeding its 4GB ceiling.
- ✗
Switch to Vertex AI Batch Prediction.
Why it's wrong here
Batch Prediction handles offline, queued inference jobs and cannot serve the low-latency online requests the endpoint exists for. The OOM stems from the container's 4GB limit against a 3GB vocabulary plus runtime overhead, so the fix is raising memory or shrinking the model footprint.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.