easyMultiple Select
PMLE Practice Question: Which TWO options are best practices for reducing…
Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)
⚠ Common exam trap
PMLE often tests whether candidates confuse throughput optimizations (bigger machines, batching) with latency optimizations — the trap is picking 'larger machine type' when the question specifically asks about latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Optimize the model using quantization or pruning
Option B is correct because quantization (e.g., reducing weights from FP32 to INT8) and pruning (removing redundant parameters) shrink the model size and reduce the compute required per inference, directly lowering prediction latency on Vertex AI Endpoints. Option C is correct because deploying the model in the same region as the clients minimizes network round-trip time, which is a significant component of end-to-end serving latency for online predictions. Option A is not a best practice for latency specifically: a larger machine with more memory may help with throughput or memory-bound models, but it does not inherently reduce per-request latency and increases cost. Option D is wrong because batch prediction is an asynchronous, offline mode that does not serve real-time requests and is not a latency optimization for online endpoints. Option E is not a supported Vertex AI Endpoint feature; there is no endpoint-level 'model caching' toggle that reduces serving latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger machine type with more memory
Why it's wrong here
Adding memory does not reduce inference latency; the bottleneck is usually compute, model size or request concurrency, and larger machines can even add cold-start delay. Bigger machine types suit memory-bound or high-throughput workloads, but latency reduction comes from optimised runtimes, accelerators or autoscaling.
- ✓
Optimize the model using quantization or pruning
Why this is correct
Quantization reduces weight precision and pruning removes redundant parameters, shrinking the model so each inference requires fewer compute cycles and less memory bandwidth. This directly lowers per-request serving latency on Vertex AI Endpoints, satisfying the stem's latency-reduction constraint without changing the endpoint's infrastructure.
- ✓
Deploy the model in the same region as the clients
Why this is correct
Co-locating the endpoint with clients in one region shortens the network round-trip distance, cutting transmission latency that dominates when model compute is already fast. This satisfies the stem's latency-reduction goal by removing geographic delay rather than altering the model itself.
- ✗
Use batch prediction instead of online prediction
Why it's wrong here
Batch prediction processes stored data asynchronously and returns results to a destination, so it does not serve interactive requests at all; latency is measured differently and often higher. Batch prediction is the right choice for large offline scoring jobs, not for reducing online endpoint response times.
- ✗
Enable model caching at the endpoint
Why it's wrong here
Vertex AI Endpoints provide no endpoint-level model caching feature; caching must be implemented in application code or a serving layer. Response caching is genuinely valuable for repeated identical queries, but it is not a configurable Vertex AI endpoint capability, so it cannot be the stated best practice.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.