Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A company needs to serve a high-throughput prediction service with strict latency requirements. They want to minimize cold starts and ensure consistent performance. Which endpoint configuration is most appropriate?

⚠ Common exam trap

Google Cloud often tests the misconception that setting min_replicas to 0 is cost-effective, but the trap here is that it ignores the strict latency requirement and the reality of cold start delays in model serving with Vertex AI.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set min_replicas to an estimated baseline and max_replicas to a higher number

Setting min_replicas to an estimated baseline ensures that a minimum number of instances are always running, eliminating cold starts for baseline traffic. Setting max_replicas to a higher number allows the service to scale up to handle traffic spikes while maintaining consistent performance. This configuration balances cost and latency by avoiding the overhead of scaling from zero while still accommodating bursts.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set min_replicas to an estimated baseline and max_replicas to a higher number

    Why this is correct

    Pinning min_replicas to an estimated baseline keeps warm capacity always available, eliminating cold starts, while max_replicas absorbs traffic spikes. This satisfies the stem's high-throughput, strict-latency and consistent-performance requirements, unlike autoscaling from zero, which reintroduces cold-start delays.

  • ✗

    Set min_replicas and max_replicas equal to a fixed number

    Why it's wrong here

    Pinning min_replicas equal to max_replicas removes autoscaling headroom, so the endpoint cannot add replicas when throughput spikes, causing request queuing that breaches latency targets. It appeals because a fixed replica count guarantees warm capacity for steady, predictable traffic.

  • ✗

    Set min_replicas to 0 and max_replicas to a high number

    Why it's wrong here

    Setting min_replicas to 0 permits the endpoint to scale down to zero, so cold starts occur whenever traffic resumes, breaking the strict latency requirement. It tempts because scaling to zero minimises cost for sporadic, latency-tolerant prediction workloads where occasional startup delay is acceptable.

  • ✗

    Do not set min_replicas; let Vertex AI automatically determine

    Why it's wrong here

    Omitting min_replicas lets Vertex AI scale to zero during idle periods, so the first request after a lull triggers a cold start and violates the strict latency requirement. It appeals because automatic replica management suits variable, latency-tolerant workloads where cost matters more than consistent response times.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

2 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. You have a Vertex AI endpoint with autoscaling enabled. You notice that during traffic spikes, the endpoint takes a long time to scale up, causing prediction errors. What is the most effective solution?

medium
  • A.Use a larger machine type to handle more requests per replica.
  • B.Reduce the target CPU utilization to trigger scaling earlier.
  • C.Increase the maximum number of replicas.
  • ✓ D.Increase the minimum number of replicas to maintain a larger buffer.

Why D: Increasing the minimum number of replicas ensures that a baseline capacity is always running, so when traffic spikes occur, the endpoint already has enough replicas to handle the initial surge without waiting for autoscaling to add more. Autoscaling in Vertex AI adds replicas based on metrics like CPU utilization, but there is inherent latency in detecting the spike, provisioning new replicas, and warming them up. By maintaining a larger buffer of minimum replicas, you reduce the reliance on rapid scale-up during spikes, thereby minimizing prediction errors. This is the most effective solution because it directly addresses the cold-start delay of autoscaling.

Variation 2. A model deployed on Vertex AI Endpoints shows increasing prediction latency. What is the most scalable way to reduce latency?

easy
  • A.Switch to a larger machine type
  • ✓ B.Enable autoscaling with min nodes increased
  • C.Use batch prediction instead
  • D.Deploy multiple model versions

Why B: Increasing the minimum number of nodes in autoscaling ensures that a baseline of compute capacity is always ready to handle requests, reducing cold-start latency. This is the most scalable approach because it allows the endpoint to dynamically scale up during traffic spikes while maintaining a floor of pre-warmed instances, directly addressing prediction latency without over-provisioning.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.