Google PCA Practice Question: Analyze and optimize technical and business processes
A retail company runs a customer-facing API on GKE Autopilot in a single region. During a quarterly sales event, traffic triples for six hours and then returns to baseline. The SRE team wants to keep the API responsive during the spike, control spend, and avoid manual intervention. They have already configured a Horizontal Pod Autoscaler based on CPU utilization with a target of 60%. Which additional action best addresses the remaining scaling bottleneck?
⚠ Common exam trap
The trap here is treating CPU utilization as a universal autoscaling signal, when request-driven workloads often saturate on concurrency or downstream latency before CPU becomes the limiting factor.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Expose a custom metric such as requests per second and configure the HPA to scale on it.
CPU-based autoscaling reacts after CPU rises, which is too late for a sudden threefold traffic surge, particularly when the API spends time waiting on network or database calls. Scaling on a request-oriented custom metric tied to actual load lets the HPA add pods proactively. This keeps latency stable during the event and lets replicas fall back to baseline afterward, satisfying responsiveness, cost control, and automation goals together.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add a PodDisruptionBudget and increase the minimum replica count in the Deployment.
Why it's wrong here
A PodDisruptionBudget protects availability during voluntary disruptions such as node drains, and a higher minimum replica count raises the floor. Neither responds to real-time traffic growth, so the API could still saturate during the six-hour spike. These settings affect resilience and baseline capacity, not dynamic scale-out.
- ✓
Expose a custom metric such as requests per second and configure the HPA to scale on it.
Why this is correct
CPU utilization lags behind actual request load, especially for I/O-bound APIs that block on downstream calls while CPU stays moderate. Scaling on a request-rate metric lets the HPA add replicas as soon as traffic climbs, matching the sales-event pattern. This reduces latency risk and avoids over-provisioning between events, directly addressing the bottleneck.
- ✗
Configure a Vertical Pod Autoscaler in recommendation mode to right-size pod requests.
Why it's wrong here
VPA recommendations help set accurate CPU and memory requests over time, but they do not add replicas during a traffic spike. In recommendation-only mode it changes nothing at runtime. Even in auto mode, VPA restarts pods to apply new sizes, which is disruptive during a sales event and does not provide horizontal elasticity.
- ✗
Increase the HPA target CPU utilization to 80% so pods are added more aggressively.
Why it's wrong here
Raising the CPU target makes the HPA less sensitive, not more, because it waits for higher utilization before adding replicas. During a sharp threefold spike this delays scale-out and can cause latency SLO violations. It also does not address the underlying issue that CPU is a lagging indicator of request load.
Go deeper
Related to this question
Learn chapter
Security Best Practices and Compliance
Key term
Pod
A pod is the smallest deployable unit in Kubernetes, containing one or more containers that share storage, network, and a specification for how to run.
Key term
GKE Autopilot
GKE Autopilot is a managed mode of Google Kubernetes Engine that automatically handles node provisioning, scaling, and maintenance so you only pay for your running pods.
About these practice questions
This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.