PMLE Serving and Scaling Models Practice Question
A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)
⚠ Common exam trap
A common misconception is that scale-to-zero (min_replicas = 0) is always cost-effective, but in latency-sensitive real-time inference, it introduces unacceptable cold-start delays, making baseline warm instances (min_replicas > 0) essential.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.
Option B is correct because a GPU-enabled machine type such as an N1 instance with an NVIDIA T4 accelerator provides the parallel compute throughput needed to keep TensorFlow inference within a p99 latency SLO under 100ms, which CPU-only serving often cannot guarantee for larger models. Option C is correct because setting min_replicas = 3 keeps a baseline of warm, already-loaded model instances, eliminating cold-start latency for the initial requests and giving the autoscaler headroom to absorb traffic spikes without waiting for new replicas to initialize. Option A is not appropriate because min_replicas = 0 enables scale-to-zero, which directly reintroduces cold-start latency and violates the strict p99 < 100ms SLO. Option D is not selected because Vertex AI Model Optimization (quantization) is a model-compression technique that may reduce latency but is not a required configuration for meeting the SLO and can degrade accuracy. Option E is not selected because batch prediction processes data offline in bulk and cannot serve real-time fraud detection requests with sub-100ms latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set min_replicas = 0 to allow scale-to-zero and save costs.
Why it's wrong here
Scale-to-zero removes all warm instances, so every spike request triggers a cold start, breaching the p99 < 100ms SLO. Keeping min_replicas above zero preserves warm capacity. Scale-to-zero suits batch or latency-tolerant workloads where idle cost saving outweighs startup delay.
- ✓
Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.
Why this is correct
GPU acceleration (N1 with T4) cuts inference compute time, directly addressing the p99 < 100ms SLO that CPU inference on a TensorFlow fraud model would likely breach. It does not solve cold starts, so it pairs with warm replicas.
- ✓
Set min_replicas = 3 to keep a baseline of warm instances.
Why this is correct
Setting min_replicas = 3 keeps three warm instances permanently provisioned, so incoming requests hit already-loaded containers rather than triggering a fresh model load. This directly satisfies the stem's cold-start minimisation requirement, since the p99 < 100ms SLO cannot absorb TensorFlow initialisation delays during traffic spikes.
- ✗
Enable Vertex AI Model Optimization for automatic quantization.
Why it's wrong here
Quantisation reduces model size and inference cost but does not by itself minimise cold starts; that requires minimum replica counts or pre-warmed instances. Model Optimization suits throughput and memory-constrained serving, not the cold-start requirement in this scenario.
- ✗
Use batch prediction instead of online prediction.
Why it's wrong here
Batch prediction processes accumulated data offline, returning results after the job completes, so it cannot satisfy a p99 under 100ms per request. It suits scheduled scoring of large stored datasets; real-time fraud detection with strict latency demands online prediction endpoints.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.