Google PCA Design and plan a cloud solution architecture Practice Question
A media company is preparing to migrate a batch reporting application to Google Cloud. The application currently runs on physical servers that are used at about 20 percent CPU on average, but it has two short month-end peaks each quarter when utilization reaches 90 percent for about six hours. The company wants to reduce infrastructure cost while guaranteeing the application always has enough capacity during the peaks. What should the architect recommend?
⚠ Common exam trap
The trap here is equating a committed use discount or a high autoscaler minimum with cost optimization, when both lock in peak capacity that runs during the long low-utilization periods.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the application on a managed instance group with autoscaling based on CPU utilization, with a minimum size that covers baseline load and a maximum size that covers the peaks.
The workload is spiky, with a low average and short, predictable peaks. Horizontal autoscaling on managed instance groups matches that shape: the group grows during the month-end surge and shrinks afterward, so cost tracks actual demand while the maximum size preserves guaranteed headroom. Fixed peak sizing, cluster autoscaler with a peak-sized minimum, and pinned Cloud Run minimum instances all keep peak capacity running continuously.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Deploy the application on a managed instance group with autoscaling based on CPU utilization, with a minimum size that covers baseline load and a maximum size that covers the peaks.
Why this is correct
An autoscaling managed instance group adds VMs when CPU rises and removes them when demand falls, so the company pays for baseline capacity most of the time and scales out only during the month-end peaks. Setting the minimum to baseline and the maximum to peak capacity guarantees headroom. This matches the workload's spiky profile directly.
- ✗
Deploy the application on a managed instance group of Compute Engine VMs sized for the month-end peak, using a committed use discount for the full capacity.
Why it's wrong here
Sizing permanently for the peak and committing to it guarantees capacity but pays for peak capacity around the clock, which is exactly the waste the company wants to eliminate. Committed use discounts lower the per-unit price of that reserved capacity but do not reduce how much capacity you hold. Average utilization would sit near 20 percent while you pay for 100 percent.
- ✗
Deploy the application to a Google Kubernetes Engine cluster with a single large node pool and enable cluster autoscaler with a minimum node count equal to the peak requirement.
Why it's wrong here
Enabling cluster autoscaler is directionally right, but setting the minimum node count to the peak requirement means those nodes are always running, so the cluster never scales down below peak. That recreates the fixed-capacity cost problem inside a more complex platform. The minimum should cover baseline load, and the maximum should cover the peaks.
- ✗
Deploy the application to Cloud Run with a minimum instance count set to the number of instances needed at peak and concurrency set to one.
Why it's wrong here
Cloud Run is a strong fit for request-driven services, but a batch reporting application is not necessarily containerized or request-triggered, and pinning minimum instances to the peak count again pays for peak capacity continuously. Setting concurrency to one also wastes per-instance capacity. This neither guarantees the right capacity model nor reduces cost.
Go deeper
Related to this question
Learn chapter
Cloud SQL and Managed Data Stores
Key term
Batch
Batch is a cloud computing service that runs large numbers of computing jobs as a group, or batch, without needing to manage individual servers.
Key term
Autoscaler
An Autoscaler is a cloud service that automatically increases or decreases the number of virtual machines (instances) or resources based on real-time demand, so your application always has enough capacity without wasting money on idle servers.
About these practice questions
Courseiva writes every PCA question from scratch — 807 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.