Courseiva

PMLE Serving and Scaling Models Practice Question

A fintech company needs to deploy a TensorFlow model for real-time fraud detection with strict latency SLO (p99 < 100ms). They expect variable traffic with spikes. They also want to minimize cold-start latency. Which two configurations should they use? (Choose 2)

⚠ Common exam trap

A common misconception is that scale-to-zero (min_replicas = 0) is always cost-effective, but in latency-sensitive real-time inference, it introduces unacceptable cold-start delays, making baseline warm instances (min_replicas > 0) essential.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.

Option B is correct because a GPU-enabled machine type such as an N1 instance with an NVIDIA T4 accelerator provides the parallel compute throughput needed to keep TensorFlow inference within a p99 latency SLO under 100ms, which CPU-only serving often cannot guarantee for larger models. Option C is correct because setting min_replicas = 3 keeps a baseline of warm, already-loaded model instances, eliminating cold-start latency for the initial requests and giving the autoscaler headroom to absorb traffic spikes without waiting for new replicas to initialize. Option A is not appropriate because min_replicas = 0 enables scale-to-zero, which directly reintroduces cold-start latency and violates the strict p99 < 100ms SLO. Option D is not selected because Vertex AI Model Optimization (quantization) is a model-compression technique that may reduce latency but is not a required configuration for meeting the SLO and can degrade accuracy. Option E is not selected because batch prediction processes data offline in bulk and cannot serve real-time fraud detection requests with sub-100ms latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set min_replicas = 0 to allow scale-to-zero and save costs.

    Why it's wrong here

    Scale-to-zero removes all warm instances, so every spike request triggers a cold start, breaching the p99 < 100ms SLO. Keeping min_replicas above zero preserves warm capacity. Scale-to-zero suits batch or latency-tolerant workloads where idle cost saving outweighs startup delay.

  • ✓

    Use a GPU-enabled machine type (e.g., N1 with T4) to accelerate inference.

    Why this is correct

    GPU acceleration (N1 with T4) cuts inference compute time, directly addressing the p99 < 100ms SLO that CPU inference on a TensorFlow fraud model would likely breach. It does not solve cold starts, so it pairs with warm replicas.

  • ✓

    Set min_replicas = 3 to keep a baseline of warm instances.

    Why this is correct

    Setting min_replicas = 3 keeps three warm instances permanently provisioned, so incoming requests hit already-loaded containers rather than triggering a fresh model load. This directly satisfies the stem's cold-start minimisation requirement, since the p99 < 100ms SLO cannot absorb TensorFlow initialisation delays during traffic spikes.

  • ✗

    Enable Vertex AI Model Optimization for automatic quantization.

    Why it's wrong here

    Quantisation reduces model size and inference cost but does not by itself minimise cold starts; that requires minimum replica counts or pre-warmed instances. Model Optimization suits throughput and memory-constrained serving, not the cold-start requirement in this scenario.

  • ✗

    Use batch prediction instead of online prediction.

    Why it's wrong here

    Batch prediction processes accumulated data offline, returning results after the job completes, so it cannot satisfy a p99 under 100ms per request. It suits scheduled scoring of large stored datasets; real-time fraud detection with strict latency demands online prediction endpoints.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.