Courseiva

PMLE Serving and Scaling Models Practice Question

You have a Vertex AI endpoint serving a model that returns predictions in about 200 ms. During a marketing campaign, traffic increases tenfold for short bursts. You want the endpoint to handle the bursts without manual intervention and without over-provisioning for the entire day. What should you do?

⚠ Common exam trap

The trap here is thinking that a larger machine type provides elasticity, when autoscaling adds replicas horizontally and is what actually handles burst traffic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure autoscaling on the endpoint with a minimum replica count of one and a maximum replica count that covers peak load.

Vertex AI endpoints support autoscaling based on metrics such as CPU utilization or request concurrency. By setting a low minimum replica count, you keep baseline cost low during normal periods. A maximum replica count sized for peak load allows the endpoint to scale out automatically when traffic increases tenfold. This elastic behavior handles short bursts without manual intervention and avoids paying for peak capacity all day, which is exactly what the scenario requires.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the machine type to a larger instance with more memory and vCPUs.

    Why it's wrong here

    A larger machine type increases per-replica capacity but does not automatically add replicas during bursts. If traffic increases tenfold, a single larger replica may still be overwhelmed, and there is no horizontal scaling. This approach also over-provisions resources during normal hours, increasing cost without providing the elastic response needed for short, intense spikes.

  • ✓

    Configure autoscaling on the endpoint with a minimum replica count of one and a maximum replica count that covers peak load.

    Why this is correct

    Vertex AI endpoint autoscaling adjusts replicas based on utilization or request concurrency. Setting a low minimum keeps cost down during normal traffic, while a maximum that covers peak load allows the endpoint to scale out during bursts. This matches the requirement to handle tenfold spikes automatically without over-provisioning all day, because replicas are added only when needed and removed when traffic subsides.

  • ✗

    Deploy the model to a batch prediction job and schedule it to run every hour.

    Why it's wrong here

    Batch prediction is designed for asynchronous, offline scoring of large datasets, not for real-time burst traffic. It cannot serve individual low-latency requests during a marketing campaign. Scheduling hourly batch jobs would introduce significant delay and would not respond to the immediate tenfold increase in interactive prediction requests. This approach fails to meet the real-time serving requirement.

  • ✗

    Enable request logging and set up an alert when latency exceeds a threshold.

    Why it's wrong here

    Logging and alerting provide observability but do not add capacity. When traffic spikes tenfold, the endpoint would still be under-provisioned, and alerts would only notify you after latency degrades. Without autoscaling or additional replicas, the endpoint cannot handle the burst automatically. This option addresses detection, not the requirement to handle the load without manual intervention.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.