Courseiva
hardMultiple Choice

PMLE Practice Question: A data engineer is troubleshooting a Vertex AI…

A data engineer is troubleshooting a Vertex AI Endpoint that serves a large BERT model. After deployment, many prediction requests fail with 'Out of Memory' errors. The machine type is n1-standard-8 (30 GB memory) with no accelerator. Which action will most likely resolve the issue?

⚠ Common exam trap

PMLE often tests the misconception that adding a GPU fixes OOM errors, when OOM is a host memory issue and the correct fix is a high-memory machine type.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Change the machine type to n1-highmem-16 (104 GB memory).

The Out of Memory errors occur because the n1-standard-8 has only 30 GB RAM, which is insufficient to load the large BERT model plus runtime overhead. Switching to n1-highmem-16 provides 104 GB RAM, giving ample headroom for the model and inference batch, directly resolving the OOM condition.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Change the machine type to n1-highmem-16 (104 GB memory).

    Why this is correct

    The n1-standard-8's 30 GB cannot hold the large BERT model plus inference tensors, causing out-of-memory failures. Moving to n1-highmem-16 raises memory to 104 GB, satisfying the model's footprint without accelerators. More vCPUs alone would not increase memory proportionally.

  • ✗

    Use batch prediction instead of online prediction.

    Why it's wrong here

    Batch prediction processes inputs asynchronously from Cloud Storage and returns results to a bucket, so it cannot serve the live endpoint requests failing with out-of-memory errors. It is correct for large offline scoring jobs, not for online serving that requires immediate responses.

  • ✗

    Add a GPU accelerator (e.g., NVIDIA T4) to offload computation.

    Why it's wrong here

    GPU helps compute but does not increase available RAM.

  • ✗

    Quantize the model from FP32 to INT8.

    Why it's wrong here

    Quantisation shrinks weights and activations, cutting memory use, but it changes numerical precision and requires re-exporting or re-converting the model, which the stem does not mention as available. It is tempting because INT8 is the standard remedy for memory-bound inference, and would be correct if the model could be re-quantised and accuracy loss were acceptable.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.