Question 441 of 486
GKE Cost Optimization: Preemptible VMs and Cluster Autoscaler
A company runs a critical application on Google Kubernetes Engine (GKE) with 3 nodes (e2-standard-4). To reduce costs, the team is considering right-sizing the nodes. The application is latency-sensitive and experiences periodic traffic spikes. What is the most cost-effective approach that maintains performance during spikes?
Quick Answer
The answer is to create a node pool with smaller e2-standard-2 nodes using preemptible VMs and enable the cluster autoscaler. This is correct because preemptible VMs offer a substantial discount—up to 60-91%—over regular instances, making them ideal for reducing baseline costs, while the cluster autoscaler dynamically adds nodes during traffic spikes to maintain performance for latency-sensitive workloads. On the Google Professional Cloud DevOps Engineer exam, this scenario tests your understanding of GKE cost optimization strategies, specifically how to balance preemptible VMs with scaling policies; a common trap is choosing larger static nodes, which wastes resources during low traffic. Memory tip: think “small and cheap for the base, autoscale for the burst.”
⚠ Common exam trap
It's easy for candidates to assume preemptible VMs are unsuitable for production or latency-sensitive workloads, but the cluster autoscaler mitigates the risk of VM termination by quickly replacing nodes, making this a valid cost-saving strategy for spike-tolerant applications.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a node pool with smaller nodes (e2-standard-2) using preemptible VMs and enable cluster autoscaler.
The most cost-effective approach because it uses smaller e2-standard-2 nodes with preemptible VMs, which are significantly cheaper than regular VMs, and combines this with the cluster autoscaler to automatically add nodes during traffic spikes. This ensures that the latency-sensitive application maintains performance by scaling out horizontally when needed, while minimizing baseline costs with smaller, cheaper nodes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a single larger node (n2-standard-8) to reduce node count and network overhead.
Why it's wrong here
This creates a single point of failure and may not handle spikes efficiently.
- ✗
Switch to 3 n2-standard-2 nodes to reduce vCPU and memory, and rely on horizontal pod autoscaling.
Why it's wrong here
This may cause resource contention during spikes as the smaller nodes have less capacity.
- ✓
Create a node pool with smaller nodes (e2-standard-2) using preemptible VMs and enable cluster autoscaler.
Why this is correct
Preemptible VMs are cheaper, and autoscaler adds nodes during spikes, maintaining performance.
- ✗
Keep current node size but use committed use discounts for 1 year to reduce per-hour cost.
Why it's wrong here
Committed use discounts require upfront commitment and may not be optimal if right-sizing is possible.
About these practice questions
Courseiva creates original exam-style practice questions with explanations and wrong-answer analysis. It does not publish real exam questions, exam dumps, or protected exam content. Learn why practice questions differ from exam dumps →
Same concept, more angles
2 more ways this is tested on PCDOE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A company runs a web application on Google Kubernetes Engine (GKE) with multiple services. They want to reduce costs without impacting performance. Which THREE actions should they take? (Choose three.)
hard- A.Enable cluster autoscaling and manually scale nodes based on peak load.
- B.Deploy a service mesh like Istio to optimize traffic routing.
- ✓ C.Enable node auto-provisioning to automatically adjust node pools.
- ✓ D.Right-size CPU and memory requests and limits for each service.
- ✓ E.Use preemptible VMs for stateless, fault-tolerant workloads.
Why C: Node auto-provisioning in GKE automatically creates and scales node pools based on the resource requirements of pending pods. This eliminates the need for manual node pool management and ensures that only the necessary compute resources are provisioned, reducing costs without manual intervention or over-provisioning.
Variation 2. A company is using Google Kubernetes Engine (GKE) with multiple node pools. They notice that their monthly costs are higher than expected. Upon review, they find that several preemptible VMs are being recreated frequently, leading to sustained usage costs. What is the most cost-effective solution to reduce costs?
medium- A.Purchase committed use discounts for the preemptible VMs.
- B.Increase the number of preemptible VMs to spread the workload.
- C.Enable sustained use discounts for the existing VMs.
- ✓ D.Migrate to Spot VMs, which have a lower price and no maximum runtime.
Why D: Spot VMs are the recommended replacement for preemptible VMs, offering the same low price but without the 24-hour maximum runtime limit. This eliminates the frequent recreation and sustained usage costs caused by preemptible VMs being terminated and restarted, directly reducing monthly expenses.
Last reviewed: Jun 11, 2026
This PCDOE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCDOE exam.
Question Discussion
Share a tip, memory trick, or ask about the reasoning behind this question. Do not post real exam questions, leaked content, braindumps, or copyrighted exam material. Comments are moderated and may be removed without notice.
Sign in to join the discussion.