easyMultiple Choice
Configure Autoscaling and GPU for Traffic Spikes
You need to serve a TensorFlow model that has a cold start latency of 20 seconds. The model is used for a real-time application with unpredictable traffic, but occasional bursts require immediate responses. What is the best deployment strategy to minimize both cold start impact and cost?
⚠ Common exam trap
The trap is assuming a bigger machine or serverless platform eliminates cold start, when the real lever is keeping at least one replica warm while letting autoscaling handle bursts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set min_replica_count to 1 to keep at least one instance always warm.
Setting min_replica_count to 1 keeps one model instance always loaded, eliminating the 20-second cold start for the first request and ensuring immediate responses during bursts. Because only one replica is kept warm, cost stays low compared to maintaining multiple always-on replicas, and autoscaling can add more when traffic spikes. This balances latency and cost for unpredictable real-time workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set min_replica_count to 1 to keep at least one instance always warm.
Why this is correct
Setting min_replica_count to 1 keeps one replica permanently loaded, so the 20-second TensorFlow cold start is paid only at deployment rather than on each burst. Bursts are absorbed by scaling additional replicas, and idle capacity stays minimal, balancing latency against cost.
- ✗
Use a larger machine type to reduce cold start time.
Why it's wrong here
A larger machine type shortens initialisation marginally but does not eliminate the 20-second TensorFlow load, and it raises cost continuously while sitting idle between unpredictable bursts. It is tempting because bigger instances genuinely reduce some framework startup overhead, which would help if the cold start were compute-bound rather than model-load-bound.
- ✗
Set min_replica_count to 0 and rely on autoscaling to handle bursts.
Why it's wrong here
Setting min_replica_count to 0 means every burst triggers a fresh 20-second cold start, directly violating the immediate-response requirement. It is tempting because scale-to-zero minimises idle cost, which would be correct for batch or latency-tolerant workloads where occasional slow first requests are acceptable.
- ✗
Enable serving on Cloud Run for faster cold start.
Why it's wrong here
Cloud Run's request-based cold start still incurs the model's 20-second load on scale-out, so bursts are not served immediately. It is tempting because Cloud Run scales to zero and bills per request, which would be the right choice for latency-tolerant or infrequent inference where cost outweighs instant response.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company is serving a model for their e-commerce website. They expect traffic to be low at night and very high during flash sales. They want to minimize costs while ensuring availability during spikes. Which autoscaling configuration should they use?
easy- A.min_replica_count=5, max_replica_count=5, target_cpu=60
- ✓ B.min_replica_count=1, max_replica_count=20, target_cpu=60
- C.min_replica_count=10, max_replica_count=10, target_cpu=60
- D.min_replica_count=0, max_replica_count=100, target_cpu=80
Why B: Setting a high max_replica_count allows scaling to handle spikes, while a low min_replica_count saves cost during low traffic. CPU utilization target of 60% is reasonable.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.