Courseiva

Google PCA Design and plan a cloud solution architecture Practice Question

A company runs a multi-tier web application on Google Kubernetes Engine (GKE) with a frontend service, a backend service, and a Cloud SQL for PostgreSQL database. During peak hours, the frontend pod CPU usage is high (consistently above 80%), while the backend service shows moderate CPU usage (around 50%). Response times for user requests increase significantly, often exceeding the 200ms p99 latency target. Cloud SQL metrics show low query latency and no contention. The team wants to improve performance in a cost-effective manner. Which initial step should they take?

⚠ Common exam trap

Google Cloud often tests the misconception that backend or database changes are needed when the bottleneck is clearly at the frontend tier, leading candidates to choose expensive or irrelevant scaling options like read replicas or vertical scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the number of frontend pods by adjusting the horizontal pod autoscaler's target CPU utilization.

The frontend pods are CPU-bound during peak hours, causing increased response times. Increasing the number of frontend pods via the Horizontal Pod Autoscaler (HPA) by lowering the target CPU utilization threshold distributes the load across more replicas, directly addressing the bottleneck without additional infrastructure cost. This is the most cost-effective initial step because it leverages existing resources and autoscaling capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Add a read replica for Cloud SQL to offload read queries.

    Why it's wrong here

    Database latency is low, so a read replica would not improve performance and would incur additional cost.

  • ✗

    Migrate the backend service to a custom machine type with more vCPUs.

    Why it's wrong here

    The backend is not the bottleneck; upgrading its resources would be expensive and unlikely to improve overall latency.

  • ✗

    Enable vertical pod autoscaling for the backend service.

    Why it's wrong here

    Backend CPU is only 50% and not the bottleneck; vertical scaling could help but is less impactful and may increase cost without addressing the frontend.

  • ✓

    Increase the number of frontend pods by adjusting the horizontal pod autoscaler's target CPU utilization.

    Why this is correct

    Frontend CPU is high, so scaling out frontend pods will help handle the load and reduce latency. This is cost-effective as it adds only needed capacity.

About these practice questions

This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.