Courseiva
hardMultiple Choice

PMLE Practice Question: A team is scaling their prototype inference model…

A team is scaling their prototype inference model to handle high-throughput requests with low latency. They use a custom container on Vertex AI Prediction. They notice that latency spikes occur under heavy load. What is the most effective strategy?

⚠ Common exam trap

PMLE often tests the reflex to throw hardware (bigger machine, GPU) at latency problems, when the expected answer is serving-layer optimization like batching and warm-up.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Optimize model serving with batching and model warm-up.

Batching groups multiple inference requests into a single forward pass, dramatically improving GPU/CPU utilization and throughput, while model warm-up pre-loads weights and initializes the runtime so the first requests after a scale-out event do not pay cold-start latency. Together they address the latency spikes under heavy load without over-provisioning. This is the standard Vertex AI Prediction optimization pattern for custom containers.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable auto-scaling with a higher minimum number of replicas.

    Why it's wrong here

    A higher minimum replica count only pre-warms capacity; it does not scale out during the load spike, so latency still climbs once those replicas saturate. Auto-scaling with a higher minimum suits steady baseline traffic, whereas the spike demands metric-driven scale-out on the serving replica count.

  • ✓

    Optimize model serving with batching and model warm-up.

    Why this is correct

    Batching amortises per-request overhead across concurrent inputs, while warm-up pre-loads weights and initialises the container so the first requests avoid cold-start latency. Together they directly address the latency spikes observed under heavy load on the custom container.

  • ✗

    Use a larger machine type with more CPUs.

    Why it's wrong here

    Adding CPUs cannot reduce queueing latency once the container's request-handling threads saturate; Vertex AI scales throughput by adding replicas, not by enlarging a single machine. Larger machine types suit memory-bound or GPU-bound single-replica workloads, not horizontal high-throughput serving.

  • ✗

    Use a GPU-based machine.

    Why it's wrong here

    May improve throughput but not necessarily tail latency spikes.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.