Courseiva

PMLE Serving and Scaling Models Practice Question

You are deploying a model to a Vertex AI endpoint and need to minimize latency for online predictions. Which machine type should you choose?

⚠ Common exam trap

The trap is overlooking the need for a GPU for low-latency inference and selecting a CPU-only machine type based on cost or memory, when the question explicitly asks to minimize latency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

n1-standard-2 with NVIDIA Tesla T4

For online predictions with minimal latency, a machine type with a GPU is essential because it accelerates model inference. The n1-standard-2 with NVIDIA Tesla T4 provides GPU acceleration, which significantly reduces latency compared to CPU-only instances. The other options lack GPUs and are not optimized for low-latency inference.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    n1-standard-2 with NVIDIA Tesla T4

    Why this is correct

    NVIDIA Tesla T4 GPUs deliver low-latency inference for online predictions, and the n1-standard-2 shape supplies adequate vCPU and memory for a single model replica. This satisfies the stem's latency-minimisation constraint, since T4s accelerate matrix operations that CPU-only machine types would process far more slowly.

  • ✗

    e2-standard-2

    Why it's wrong here

    e2-standard-2 is a cost-optimised shared-core machine type with variable CPU performance, so tail latency is unpredictable under load. It is tempting for low-cost, non-latency-critical serving or batch inference, but online prediction requires the consistent, dedicated vCPU performance that e2 does not guarantee.

  • ✗

    n1-standard-2

    Why it's wrong here

    n1-standard-2 provides only 2 vCPUs and 7.5 GB RAM, capping parallel inference throughput and raising per-request latency. It is tempting as a general-purpose balanced shape for light serving workloads, but it lacks the compute headroom that latency-sensitive online prediction demands.

  • ✗

    n1-highmem-2

    Why it's wrong here

    n1-highmem-2 allocates 13 GB RAM but only 2 vCPUs, so inference compute is throttled despite generous memory. It is tempting because high-memory shapes suit large in-memory models or data-heavy preprocessing, but online latency depends on vCPU throughput, which this shape lacks.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.