PMLE Serving and Scaling Models Practice Question
Which API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints?
⚠ Common exam trap
Google often tests the misconception that REST API is the default or only way to interact with cloud services, but the trap here is that for high-throughput, low-latency online predictions, gRPC is explicitly recommended over REST due to its performance advantages with Protocol Buffers and HTTP/2.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
gRPC API
gRPC API is recommended for high-throughput, low-latency online prediction requests to Vertex AI endpoints because it uses HTTP/2 for multiplexed streaming, binary serialization (Protocol Buffers), and supports bidirectional streaming, which reduces latency and improves throughput compared to REST. Vertex AI's prediction service natively supports gRPC for real-time inference, making it the optimal choice for latency-sensitive applications.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions is a serverless compute platform for running event-driven code, not a client interface for invoking Vertex AI endpoints. It would be correct for gluing services together, such as triggering a prediction when a file lands in Cloud Storage, but it adds a hop rather than serving low-latency requests directly.
- ✗
REST API
Why it's wrong here
REST API uses synchronous HTTP request-response cycles with serialisation overhead and no persistent connection, which introduces per-call latency and limits throughput under high concurrency compared to gRPC’s multiplexed streaming over HTTP/2. It is tempting because REST is the standard for most Vertex AI operations like model deployment, batch predictions, and resource management, and would be correct for infrequent, ad-hoc requests where simplicity of integration outweighs latency requirements.
- ✗
Cloud Pub/Sub
Why it's wrong here
Pub/Sub is for asynchronous messaging, not real-time prediction.
- ✓
gRPC API
Why this is correct
gRPC API uses HTTP/2 multiplexing and binary Protocol Buffers, cutting serialisation overhead and connection latency compared with REST's JSON over HTTP/1.1. This directly satisfies the stem's high-throughput, low-latency constraint for online prediction requests to Vertex AI endpoints.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.