Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

You are deploying a model to a Vertex AI endpoint that will serve predictions for a mobile application. The application sends a single request per user action and expects a response within 100 ms. The model is small and CPU-bound. You want to minimize cost while meeting the latency requirement. Which endpoint configuration should you choose?

⚠ Common exam trap

The trap here is assuming that GPUs or a fixed number of replicas are needed for low latency, when a small CPU model can meet the SLO with autoscaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=5, and enable autoscaling based on CPU utilization.

For a small CPU-bound model with variable traffic, the optimal cost-latency trade-off is a CPU-only machine type with a low minimum replica count and autoscaling enabled. This keeps baseline cost low while allowing the endpoint to add replicas when CPU utilization increases, preserving the 100 ms latency target. GPU instances and fixed high replica counts unnecessarily increase cost without improving latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=5, and enable autoscaling based on CPU utilization.

    Why this is correct

    This configuration uses cost-effective CPU instances and allows the endpoint to scale out when CPU utilization rises, ensuring latency remains low under load. Setting minReplicaCount=1 keeps idle cost low, while maxReplicaCount=5 provides headroom. Autoscaling based on CPU is appropriate for a CPU-bound model. This balances cost and performance effectively.

  • ✗

    Use a CPU-only machine type with minReplicaCount=5 and maxReplicaCount=5 to ensure high availability.

    Why it's wrong here

    Running five replicas at all times is expensive and unnecessary for a small model that can handle significant load per replica. It over-provisions capacity and increases cost, violating the goal to minimize cost. High availability can be achieved with fewer replicas and autoscaling. This configuration is wasteful for the described workload.

  • ✗

    Use a CPU-only machine type with minReplicaCount=1 and maxReplicaCount=1.

    Why it's wrong here

    A single CPU replica may meet latency for low traffic, but it cannot scale if traffic increases, leading to queueing and potential SLO violations. While it minimizes cost, it does not provide the elasticity needed for a mobile app that may have variable load. The requirement is to minimize cost while meeting latency, which implies the ability to handle load changes without over-provisioning.

  • ✗

    Use a GPU-enabled machine type with minReplicaCount=1 and maxReplicaCount=1.

    Why it's wrong here

    GPU instances are more expensive than CPU instances and provide no benefit for a small CPU-bound model. They also incur higher cold-start times. Since the model is small and CPU-bound, a GPU would be underutilized and increase cost without improving latency. The single replica also cannot scale, but the primary issue is the unnecessary GPU expense.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.