PMLE Serving and Scaling Models Practice Question
A data scientist wants to deploy a trained TensorFlow model to Vertex AI for online predictions. They need to serve predictions with low latency and want to leverage GPU acceleration. Which machine type should they select when creating the Vertex AI endpoint?
⚠ Common exam trap
Test-takers frequently assume any machine type can be paired with a GPU, but only specific series (like n1, n2, g2) support GPU attachment, and the e2 series explicitly does not, leading to a wrong selection if the GPU requirement is overlooked.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
n1-standard-4 with 1 NVIDIA Tesla T4
The n1-standard-4 machine type supports attaching GPUs such as the NVIDIA Tesla T4, which provides GPU acceleration for low-latency online predictions. Vertex AI endpoints require a machine type that allows GPU attachment, and the n1-series is one of the few families that supports GPUs, while the T4 offers a good balance of cost and performance for inference workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
n1-standard-4 with 1 NVIDIA Tesla T4
Why this is correct
NVIDIA Tesla T4 GPUs attached to n1-standard-4 instances provide the GPU acceleration the stem demands, while n1-standard-4 supplies sufficient vCPU and memory for low-latency online inference of a TensorFlow model. Vertex AI supports this accelerator pairing directly, satisfying both the GPU and latency constraints when deploying the endpoint.
- ✗
n1-standard-4
Why it's wrong here
n1-standard-4 is a general-purpose CPU-only machine type, so it cannot attach the GPU accelerators the scenario requires for low-latency inference. It would be the right pick for lightweight CPU-bound workloads, but GPU-backed types such as n1-standard-4 with an attached accelerator are needed here.
- ✗
e2-standard-4
Why it's wrong here
e2-standard-4 is a general-purpose machine type that cannot attach GPUs, so it delivers no GPU acceleration for inference. It is tempting because it is cost-effective for lightweight CPU serving, and it would be correct for low-throughput online prediction where latency targets are met without hardware acceleration.
- ✗
n1-highmem-8
Why it's wrong here
n1-highmem-8 is a general-purpose machine type with high memory per vCPU but no attached GPU; GPU acceleration requires an accelerator-optimised type such as n1-standard-4 with an added GPU. It is tempting for memory-hungry models, and it would be correct when serving large in-memory workloads without GPU acceleration.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.