Courseiva
easyMultiple Choice

PMLE Practice Question: Serve a model for real-time predictions with a…

A company needs to serve a model for real-time predictions with a strict latency SLA of 100ms at the 99th percentile. The model is lightweight and traffic patterns are highly variable with occasional spikes. Which deployment strategy best meets the SLA while controlling cost?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy to Vertex AI Endpoint with min_replica_count=3 and autoscaling enabled.

Setting a minimum number of replicas ensures baseline capacity to handle initial spikes without cold start delays, while autoscaling handles larger spikes. Option A is wrong because Cloud Run with autoscaling to zero may cause cold start delays, which could violate the strict latency SLA. Option B is wrong because manual scaling with a fixed number of replicas may lead to over-provisioning or under-provisioning. Option C is wrong because batch prediction is not real-time.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model as a Cloud Run service with autoscaling to zero.

    Why it's wrong here

    Scaling to zero introduces cold starts, so the first request after an idle period exceeds the 100ms 99th-percentile SLA; variable spikes worsen this. It is tempting because it minimises cost during idle periods, but its correct scenario is latency-tolerant or non-interactive workloads, not strict real-time SLAs.

  • ✗

    Deploy to Vertex AI Endpoint with manual scaling and a fixed number of replicas.

    Why it's wrong here

    Manual scaling with a fixed replica count cannot absorb occasional traffic spikes, so queued requests breach the 100ms 99th-percentile SLA, while idle replicas waste cost. It is tempting because it gives predictable capacity, but its correct scenario is steady, well-forecast load rather than highly variable traffic.

  • ✗

    Use Vertex AI Batch Prediction.

    Why it's wrong here

    Batch Prediction processes jobs asynchronously and returns results to storage, so predictions are not served per request and cannot meet a 100ms 99th-percentile latency SLA. It is tempting because it is cost-effective for large volumes, but its correct scenario is offline scoring of accumulated data, not real-time serving.

  • ✓

    Deploy to Vertex AI Endpoint with min_replica_count=3 and autoscaling enabled.

    Why this is correct

    Setting min_replica_count=3 keeps three replicas warm, eliminating cold-start latency that would breach the 100ms p99 SLA during traffic spikes. Autoscaling then absorbs variable load beyond that floor, so you pay for three always-on nodes rather than provisioning peak capacity permanently.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.