Cloud Digital Leader Scaling with Google Cloud operations Practice Question
A company runs an e-commerce platform on Google Kubernetes Engine (GKE) using autoscaling. They have a baseline workload and occasional traffic spikes during promotions. They configured a Horizontal Pod Autoscaler (HPA) for their web application pods and a Cluster Autoscaler for the node pool. The HPA targets 70% CPU utilization. During a recent sales event, traffic exceeded expectations. The operations team observed that the HPA increased the desired number of replicas to 50, but only 20 pods were running. The remaining 30 pods were in 'Pending' status. The Cluster Autoscaler logs show repeated messages: 'no capacity to scale up node pool'. The node pool is configured with a maximum of 10 nodes, each with 4 vCPUs, and currently all 10 nodes are running. The team checked the node pool's current utilization and found that nodes are near capacity. What should the team do to ensure the application scales correctly during future events?
⚠ Common exam trap
Google Cloud often tests the misconception that adjusting HPA thresholds or pod resource requests alone can solve capacity issues, when the real bottleneck is an insufficient node pool maximum, which must be increased to allow the Cluster Autoscaler to add capacity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the maximum number of nodes in the node pool to allow more capacity.
The HPA requested 50 replicas, but only 20 could be scheduled because the existing 10 nodes are near capacity and the node pool is already at its maximum size. The Cluster Autoscaler cannot add more nodes because the node pool's maximum node limit prevents it from provisioning the needed additional capacity. Increasing the maximum number of nodes in the node pool (Option C) allows the Cluster Autoscaler to add nodes and accommodate the pending pods, enabling the HPA to scale as needed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the HPA target CPU utilization to 90% to reduce the number of replicas needed.
Why it's wrong here
Raising the HPA target CPU utilization to 90% tells the autoscaler to tolerate consistently busier pods before adding replicas, which shrinks the replica count for a given load. That directly reduces application headroom, increases per-pod CPU saturation, and can cause latency spikes or throttling during traffic bursts. It also does nothing to raise the node pool's ceiling of 10 nodes, so if all existing nodes are already at full capacity, pods will still remain Pending. This approach is a performance risk, not a scaling solution.
- ✗
Reduce the pod resource requests for CPU so that more pods can fit on existing nodes.
Why it's wrong here
Reducing the CPU resource requests on the pod specification lowers the amount of CPU that Kubernetes reserves per pod, which allows the scheduler to pack more pod replicas onto the same set of nodes. However, this is an overcommitment: pods will share fewer guaranteed CPU cycles, leading to CPU contention, throttling, and degraded response times under load. It also does not change the fact that the node pool is capped at 10 nodes and the Cluster Autoscaler has no room to add compute capacity, so it merely masks the symptom with increased risk of instability.
- ✓
Increase the maximum number of nodes in the node pool to allow more capacity.
Why this is correct
Increasing the maximum number of nodes in the node pool is the correct fix because the Cluster Autoscaler has already scaled to the current maximum of 10 nodes, yet pending pods remain. By raising that ceiling, the Cluster Autoscaler can provision additional nodes to absorb the scheduler's pending pod queue, restoring capacity while preserving the CPU headroom defined by the HPA and pod resource requests. This directly addresses the root cause—the node pool's capacity ceiling—and allows both horizontal pod autoscaling and cluster autoscaling to work as intended.
- ✗
Enable extra capacity by creating a second node pool with preemptible VMs.
Why it's wrong here
Creating a second node pool with preemptible VMs does not address the root cause—the existing node pool has reached its maximum of 10 nodes, and the Cluster Autoscaler cannot add more nodes because the node pool’s capacity ceiling is already hit. Preemptible VMs are cheaper but can be terminated at any time, making them unsuitable for handling sustained traffic spikes where pod stability is critical; they would be correct if the goal were to reduce cost for fault-tolerant batch jobs.
Go deeper
Related to this question
Learn chapter
Apigee API Management Platform
Key term
Pod
A pod is the smallest deployable unit in Kubernetes, containing one or more containers that share storage, network, and a specification for how to run.
Key term
Google Kubernetes Engine
Google Kubernetes Engine (GKE) is a managed Kubernetes service on Google Cloud that lets you deploy, scale, and manage containerized applications without having to operate the underlying cluster control plane.
About these practice questions
Courseiva writes every GCDL question from scratch — 848 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.