Courseiva
easyMultiple Select

PMLE Practice Question: Which TWO options are best practices for reducing…

Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)

⚠ Common exam trap

PMLE often tests whether candidates confuse throughput optimizations (bigger machines, batching) with latency optimizations — the trap is picking 'larger machine type' when the question specifically asks about latency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Optimize the model using quantization or pruning

Option B is correct because quantization (e.g., reducing weights from FP32 to INT8) and pruning (removing redundant parameters) shrink the model size and reduce the compute required per inference, directly lowering prediction latency on Vertex AI Endpoints. Option C is correct because deploying the model in the same region as the clients minimizes network round-trip time, which is a significant component of end-to-end serving latency for online predictions. Option A is not a best practice for latency specifically: a larger machine with more memory may help with throughput or memory-bound models, but it does not inherently reduce per-request latency and increases cost. Option D is wrong because batch prediction is an asynchronous, offline mode that does not serve real-time requests and is not a latency optimization for online endpoints. Option E is not a supported Vertex AI Endpoint feature; there is no endpoint-level 'model caching' toggle that reduces serving latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a larger machine type with more memory

    Why it's wrong here

    Adding memory does not reduce inference latency; the bottleneck is usually compute, model size or request concurrency, and larger machines can even add cold-start delay. Bigger machine types suit memory-bound or high-throughput workloads, but latency reduction comes from optimised runtimes, accelerators or autoscaling.

  • ✓

    Optimize the model using quantization or pruning

    Why this is correct

    Quantization reduces weight precision and pruning removes redundant parameters, shrinking the model so each inference requires fewer compute cycles and less memory bandwidth. This directly lowers per-request serving latency on Vertex AI Endpoints, satisfying the stem's latency-reduction constraint without changing the endpoint's infrastructure.

  • ✓

    Deploy the model in the same region as the clients

    Why this is correct

    Co-locating the endpoint with clients in one region shortens the network round-trip distance, cutting transmission latency that dominates when model compute is already fast. This satisfies the stem's latency-reduction goal by removing geographic delay rather than altering the model itself.

  • ✗

    Use batch prediction instead of online prediction

    Why it's wrong here

    Batch prediction processes stored data asynchronously and returns results to a destination, so it does not serve interactive requests at all; latency is measured differently and often higher. Batch prediction is the right choice for large offline scoring jobs, not for reducing online endpoint response times.

  • ✗

    Enable model caching at the endpoint

    Why it's wrong here

    Vertex AI Endpoints provide no endpoint-level model caching feature; caching must be implemented in application code or a serving layer. Response caching is genuinely valuable for repeated identical queries, but it is not a configurable Vertex AI endpoint capability, so it cannot be the stated best practice.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.