Courseiva
easyMultiple Choice

PMLE Practice Question: Serve a model with strict latency requirements…

A company needs to serve a model with strict latency requirements (<100ms). They are using Vertex AI Prediction with CPU. During testing, latency is 150ms. What should they do?

⚠ Common exam trap

Test-takers frequently confuse throughput optimization (batching or scaling replicas) with latency reduction, failing to recognize that GPUs directly address compute-bound latency while CPU-based solutions cannot meet strict sub-100ms requirements for complex models.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch to a GPU machine type

The model's latency of 150ms exceeds the 100ms requirement. Switching to a GPU machine type (Option D) is correct because GPUs are optimized for parallel computation, significantly reducing inference latency for many ML models, especially deep learning models, compared to CPUs. Vertex AI Prediction supports GPU machine types, and this change directly addresses the latency bottleneck without altering the model or its serving configuration.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable batching to improve throughput

    Why it's wrong here

    Batching groups requests to raise throughput, but it adds queueing delay before inference starts, pushing per-request latency further above 100ms. It tempts because batching is the standard lever for cost and throughput optimisation, yet the stem's constraint is latency, not throughput.

  • ✗

    Use a smaller machine type with more replicas

    Why it's wrong here

    More replicas of a smaller machine type add horizontal capacity, not per-request speed; each individual inference still runs on slower hardware. It tempts because scaling out is the usual remedy for overloaded CPU serving, but the stem's 150ms figure reflects single-request compute time, not queue contention.

  • ✗

    Export the model to TensorFlow Lite

    Why it's wrong here

    TensorFlow Lite targets mobile and edge deployment, and converting a model from another framework risks unsupported operators and accuracy loss. It tempts because lightweight runtimes do cut inference latency, but the stem's CPU serving latency is addressed by GPU acceleration, not by changing the runtime format.

  • ✓

    Switch to a GPU machine type

    Why this is correct

    GPUs accelerate the matrix multiplications dominating neural network inference, cutting compute time well below the 100ms threshold that CPU inference currently exceeds at 150ms. For latency-bound serving, GPU machine types provide the parallel throughput needed, satisfying the stem's strict sub-100ms constraint where CPU cannot.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.