Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A company runs a high-throughput inference service on a Vertex AI Endpoint backed by a custom container. During peak hours, the endpoint's CPU utilization rises to 85%, but the autoscaler does not add replicas until utilization exceeds 95%. The team wants the autoscaler to react earlier to keep latency low. They have already deployed the model and cannot change the model artifact. What should they do?

⚠ Common exam trap

The trap here is assuming that increasing the minimum replica count changes when the autoscaler scales, when in fact it only raises the floor of running replicas.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Update the endpoint's deployed model to set a lower autoscaling metric threshold, such as 70% CPU utilization.

The autoscaler on a Vertex AI Endpoint uses a target metric and threshold defined in the deployed model's autoscaling configuration. Adjusting the CPU utilization target to a lower value makes the system add replicas before latency degrades. Other approaches either do not change the scaling trigger or introduce manual processes that are slower and less reliable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Update the endpoint's deployed model to set a lower autoscaling metric threshold, such as 70% CPU utilization.

    Why this is correct

    Vertex AI Endpoints expose an autoscaling configuration on the DeployedModel, including the target metric and threshold. Lowering the CPU utilization target to 70% makes the autoscaler add replicas sooner, which reduces latency during ramp-up. This is a configuration change on the deployed model and does not require retraining or replacing the model artifact.

  • ✗

    Enable request-response logging on the endpoint and use Cloud Monitoring alerts to manually add replicas when CPU exceeds 70%.

    Why it's wrong here

    Manual scaling via alerts is operationally fragile and cannot respond as quickly as the built-in autoscaler. Logging and alerting provide visibility but do not automatically adjust replica count. The scenario asks for the autoscaler to react earlier, which is a configuration change, not an external monitoring workflow.

  • ✗

    Increase the endpoint's minReplicaCount to match the peak traffic level so that replicas are always available.

    Why it's wrong here

    Raising minReplicaCount only guarantees a floor of replicas; it does not change the scaling trigger. During peaks above that floor, the autoscaler would still wait until 95% utilization before scaling. This approach also increases cost during all hours, including idle periods, without addressing the late scale-up behavior.

  • ✗

    Redeploy the model with a larger machine type so that each replica handles more traffic and CPU never reaches the threshold.

    Why it's wrong here

    Using a larger machine type may reduce CPU utilization per replica, but it does not change the scaling threshold and may not eliminate the need to scale. It also increases cost per replica and requires redeployment. The core issue is the autoscaling trigger, not the per-replica capacity.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.