easyMultiple Choice
PMLE Serving a model for their e-commerce website Practice Question
A company is serving a model for their e-commerce website. They expect traffic to be low at night and very high during flash sales. They want to minimize costs while ensuring availability during spikes. Which autoscaling configuration should they use?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
min_replica_count=1, max_replica_count=20, target_cpu=60
Setting a high max_replica_count allows scaling to handle spikes, while a low min_replica_count saves cost during low traffic. CPU utilization target of 60% is reasonable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
min_replica_count=5, max_replica_count=5, target_cpu=60
Why it's wrong here
A fixed replica count of five cannot expand during flash-sale spikes, so requests queue or fail once capacity saturates, and it still pays for five replicas overnight. Fixed sizing suits steady, predictable workloads where capacity planning replaces elasticity; here the stem demands both spike headroom and low-traffic cost savings.
- ✓
min_replica_count=1, max_replica_count=20, target_cpu=60
Why this is correct
Setting min_replica_count=1 keeps costs low during quiet nights, while max_replica_count=20 allows horizontal scaling to absorb flash-sale spikes. The target_cpu=60 metric triggers additional replicas before saturation, preserving availability. This balances the stem's dual constraint: minimising idle cost while guaranteeing capacity during unpredictable demand surges.
- ✗
min_replica_count=10, max_replica_count=10, target_cpu=60
Why it's wrong here
Ten permanently running replicas waste spend during the nightly lull and still cannot scale beyond ten when flash-sale traffic exceeds that ceiling, causing dropped requests. Fixed provisioning fits stable, high-baseline services; this scenario explicitly requires scaling down at night and absorbing unpredictable spikes.
- ✗
min_replica_count=0, max_replica_count=100, target_cpu=80
Why it's wrong here
Scaling from zero introduces cold-start latency, so the first requests of a flash sale may time out before replicas initialise, and an 80% CPU target reacts too slowly for sudden surges. Scale-to-zero suits intermittent, latency-tolerant batch inference, not a customer-facing site needing immediate availability.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.