Google PCA Practice Question: Analyze and optimize technical and business processes
Your company runs a multi-tier web application on Google Kubernetes Engine (GKE). The application consists of a frontend service, a backend API service, and a PostgreSQL database deployed using a StatefulSet with persistent volumes. The backend service exposes a gRPC endpoint. Recently, the team noticed that the backend service experiences intermittent high latency and occasional timeouts. The frontend service is stateless and scales well. The backend service is CPU-bound. The database is not the bottleneck. The cluster has three nodes of type n1-standard-4. The backend service is deployed with 10 replicas, each requesting 1 CPU and 2 Gi memory. Node utilization is around 70% CPU. The team suspects the network is the issue. However, after reviewing the GKE monitoring dashboard, they see that the network bytes sent/received per second for the backend pods is well below the node's network bandwidth limit. The latency spikes seem correlated with periods of high CPU throttling on the backend pods. The backend service's gRPC requests are small (under 1 KB), and the responses are also small. The team has already optimized the application code. What should the team do to reduce latency?
⚠ Common exam trap
The trap here is that candidates may focus on network or scaling solutions (A or B) because the symptom is latency, but the monitoring data explicitly points to CPU throttling, not network saturation, making CPU request adjustment the precise fix.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the CPU request for the backend pods to 2 CPUs.
The latency spikes correlate with CPU throttling, and increasing the CPU request to 2 CPUs ensures that each backend pod receives a guaranteed CPU share, reducing throttling under load. Since the backend is CPU-bound and node utilization is 70%, the current 1 CPU request may be insufficient, causing the Kubernetes CPU manager to throttle the pods when the node's CPU is contended. This directly addresses the root cause without adding unnecessary replicas or memory.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of nodes in the cluster to reduce network contention.
Why it's wrong here
Adding nodes does not raise each pod's CPU limit, so throttling persists; the dashboard already shows network throughput far below node bandwidth, ruling out contention. Scaling node count is correct when pods sit Pending due to insufficient allocatable CPU or memory across the cluster.
- ✗
Increase the number of backend replicas to 20.
Why it's wrong here
Doubling replicas spreads the same per-pod CPU request across more pods, but each pod still hits its 1 CPU ceiling and gets throttled. Horizontal scaling suits request-bound services where aggregate throughput, not per-pod CPU throttling, limits performance.
- ✓
Increase the CPU request for the backend pods to 2 CPUs.
Why this is correct
CPU throttling, not network bandwidth, causes the latency spikes. Raising each backend pod's CPU request to 2 CPUs gives the CPU-bound gRPC service enough quota to avoid throttling, directly removing the constraint correlated with the observed timeouts.
- ✗
Increase the memory request for the backend pods to 4 Gi.
Why it's wrong here
Raising memory to 4 Gi leaves the CPU request at 1 CPU, so the pods remain throttled by their CPU limit during bursts, which is the stated cause of latency. Memory tuning is the right action when pods are OOM-killed or evicted for exceeding memory limits, not for CPU throttling.
Go deeper
Related to this question
Learn chapter
Load Balancing and Autoscaling
Key term
Pod
A pod is the smallest deployable unit in Kubernetes, containing one or more containers that share storage, network, and a specification for how to run.
Key term
Anthos
Anthos is a Google Cloud platform that lets you run applications consistently across different computing environments, like on-premises data centers and multiple public clouds.
About these practice questions
One of 807 original PCA practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.