Courseiva
mediumMultiple Choice

PMLE Practice Question: A company deploys a model on Vertex AI Endpoint…

A company deploys a model on Vertex AI Endpoint and expects high traffic spikes during promotional events. The current configuration uses manual scaling with 2 replicas. Which autoscaling configuration should they use to handle spikes while minimizing cost during normal traffic?

⚠ Common exam trap

The trap is choosing 'more replicas' or 'custom metrics' as the answer, when the exam wants you to recognize that autoscaling requires both a scaling metric and min/max bounds to balance cost and performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable basic scaling with target_cpu_utilization=0.6 and set min_replica_count=2, max_replica_count=10.

Enabling autoscaling with a CPU utilization target and setting min/max replica counts allows Vertex AI to scale replicas up during traffic spikes and down during normal traffic, balancing performance and cost. The min of 2 preserves baseline capacity while the max of 10 caps cost during spikes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Keep manual scaling but increase replicas to 10.

    Why it's wrong here

    Fixed 10 replicas cannot flex with traffic, so promotional spikes beyond that capacity fail while normal periods pay for eight idle replicas. Manual provisioning suits predictable, steady demand where capacity is known in advance and cost efficiency is not the constraint.

  • ✗

    Set min_replica_count=2 and max_replica_count=10 with no scaling metric.

    Why it's wrong here

    Without a scaling metric, Vertex AI cannot trigger replica changes, so the endpoint stays at its minimum and never reaches the maximum during spikes. Bounds alone suit scenarios where an external autoscaler or scheduled scaling drives replica counts independently.

  • ✓

    Enable basic scaling with target_cpu_utilization=0.6 and set min_replica_count=2, max_replica_count=10.

    Why this is correct

    Basic scaling with target_cpu_utilization=0.6 adds replicas automatically as CPU load rises during spikes, while min_replica_count=2 keeps baseline capacity and max_replica_count=10 caps spend. This satisfies the stem's need to absorb promotional spikes while minimising cost at normal traffic.

  • ✗

    Use custom metric scaling with a Cloud Monitoring metric for prediction latency.

    Why it's wrong here

    Prediction latency is a lagging health signal, not a demand signal; it reacts only after requests queue, so replicas scale too late for sudden spikes. Custom metrics suit workloads where CPU and request counts misrepresent load, such as GPU-bound models with steady queues.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.