Courseiva
Serving and Scaling Models →mediumMultiple Select

PMLE Serving and Scaling Models Practice Question

You are deploying a model to a Vertex AI endpoint for online predictions. You need to ensure that the endpoint can handle traffic spikes and that predictions are served with low latency. Which TWO of the following configurations should you apply? (Choose two.)

⚠ Common exam trap

The trap here is assuming that logging or larger machines are sufficient for handling spikes, when the key is to have warm replicas and proactive autoscaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure autoscaling based on CPU utilization with a target utilization that triggers scale-out early.

To handle traffic spikes and ensure low latency, you should maintain a minimum replica count greater than 1 to have warm replicas ready, and configure autoscaling with a low target utilization to trigger scale-out early. These two settings together provide both baseline capacity and proactive scaling, reducing the risk of cold starts and overload during spikes. Other options like logging or larger machines do not directly address dynamic traffic handling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model to multiple endpoints and use a load balancer to distribute traffic.

    Why it's wrong here

    Deploying to multiple endpoints adds complexity and management overhead. Vertex AI endpoints already provide load balancing across replicas. Using multiple endpoints would require external load balancing and does not integrate with Vertex AI's autoscaling. This approach is redundant and does not leverage the built-in scaling capabilities of a single endpoint, making it inefficient for handling spikes and latency.

  • ✓

    Configure autoscaling based on CPU utilization with a target utilization that triggers scale-out early.

    Why this is correct

    Autoscaling based on CPU utilization with a low target utilization triggers scale-out earlier, adding replicas before the existing ones become saturated. This proactive scaling helps handle traffic spikes by provisioning additional capacity in advance. It balances cost and performance by scaling out only when needed but doing so early enough to maintain low latency. This is a standard practice for latency-sensitive endpoints.

  • ✗

    Use a larger machine type with more CPU and memory for each replica to increase per-replica throughput.

    Why it's wrong here

    A larger machine type can increase the capacity of each replica, but it does not directly address the need to handle traffic spikes dynamically. It also increases cost. Without autoscaling, a larger machine may still be overwhelmed by a spike. While it can contribute to performance, it is not as critical as ensuring multiple replicas and proactive autoscaling for handling variable traffic.

  • ✗

    Enable request-response logging to capture detailed latency metrics for each prediction.

    Why it's wrong here

    Request-response logging is useful for monitoring and debugging, but it does not directly improve the endpoint's ability to handle traffic spikes or reduce latency. It adds overhead and storage costs. While it can help identify issues, it is not a configuration that enhances performance or scalability. Therefore, it is not one of the required configurations for handling spikes and low latency.

  • ✓

    Set a minimum replica count greater than 1 to ensure that the endpoint has warm replicas ready to serve traffic.

    Why this is correct

    Setting a minimum replica count greater than 1 ensures that there are always multiple replicas available to handle incoming requests. This reduces the chance of cold starts and provides immediate capacity during traffic spikes. Warm replicas can serve requests without the delay of loading the model, which is critical for low-latency predictions. It also provides redundancy in case a replica fails.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.