PMLE Serving and Scaling Models Practice Question
You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?
⚠ Common exam trap
The trap is overlooking the need for a GPU for low-latency inference and selecting a CPU-only machine type based on cost or memory, when the question explicitly asks to minimize latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
n1-standard-2 with NVIDIA Tesla T4
For online predictions with minimal latency, a machine type with a GPU is essential because it accelerates model inference. The n1-standard-2 with NVIDIA Tesla T4 provides GPU acceleration, which significantly reduces latency compared to CPU-only instances. The other options lack GPUs and are not optimized for low-latency inference.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
n1-standard-2 with NVIDIA Tesla T4
Why this is correct
NVIDIA Tesla T4 GPUs deliver low-latency inference for online predictions, and the n1-standard-2 shape supplies adequate vCPU and memory for a single model replica. This satisfies the stem's latency-minimisation constraint, since T4s accelerate matrix operations that CPU-only machine types would process far more slowly.
- ✗
e2-standard-2
Why it's wrong here
e2-standard-2 is a cost-optimised shared-core machine type with variable CPU performance, so tail latency is unpredictable under load. It is tempting for low-cost, non-latency-critical serving or batch inference, but online prediction requires the consistent, dedicated vCPU performance that e2 does not guarantee.
- ✗
n1-standard-2
Why it's wrong here
n1-standard-2 provides only 2 vCPUs and 7.5 GB RAM, capping parallel inference throughput and raising per-request latency. It is tempting as a general-purpose balanced shape for light serving workloads, but it lacks the compute headroom that latency-sensitive online prediction demands.
- ✗
n1-highmem-2
Why it's wrong here
n1-highmem-2 allocates 13 GB RAM but only 2 vCPUs, so inference compute is throttled despite generous memory. It is tempting because high-memory shapes suit large in-memory models or data-heavy preprocessing, but online latency depends on vCPU throughput, which this shape lacks.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.