Courseiva

Google PCA Practice Question: Managing Implementation and Ensuring Solution and Operations Reliability

Your team manages a production web application on Compute Engine behind an external Application Load Balancer. During a recent incident, the load balancer's backend service marked all instances as unhealthy because the health check endpoint returned HTTP 200 but the application was actually in a degraded state. You need Cloud Monitoring to alert the operations team when the application's error rate exceeds 5% over a 5-minute window. You also need to ensure that the alert does not fire during planned maintenance windows. Which approach should you take?

⚠ Common exam trap

The trap here is assuming that a health check endpoint returning HTTP 200 means the application is healthy and that uptime checks or health check status can measure error rate.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a log-based metric that counts HTTP 5xx responses from the load balancer logs, then create an alerting policy on that metric with a threshold of 5% error rate and configure a maintenance window for planned downtime.

The requirement is to alert on application error rate exceeding 5% over 5 minutes, with suppression during maintenance. A log-based metric from load balancer logs captures every request's status, enabling an accurate error rate calculation. Alerting policies on such metrics support threshold conditions and maintenance windows. The other options either alter health checks in a harmful way, rely on uptime checks that miss application-level errors, or use sampled trace data that is not representative.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Cloud Trace to sample requests and create an alerting policy based on the latency of traces that return errors, triggering when error traces exceed 5% of total traces.

    Why it's wrong here

    Cloud Trace is designed for latency analysis and distributed tracing, not for computing error rates from all requests. Sampling means not every request is traced, so the percentage of error traces may not reflect the true error rate. Additionally, Cloud Trace does not integrate with maintenance windows for alert suppression, and the alerting condition would be based on a sampled subset rather than actual load balancer logs.

  • ✗

    Create an uptime check that sends HTTP requests to the application's health endpoint every minute and alerts when the check fails from more than one region.

    Why it's wrong here

    Uptime checks verify reachability and response codes from external locations, but they do not measure the proportion of user requests that result in errors. A degraded application might still return HTTP 200 on the health endpoint, so uptime checks would not detect the elevated error rate. This approach also lacks a mechanism to suppress alerts during maintenance windows.

  • ✗

    Configure a custom health check on the load balancer that returns HTTP 500 when the application is degraded, and create an alerting policy on the backend service's unhealthy instance count.

    Why it's wrong here

    Changing the health check to return HTTP 500 would cause the load balancer to remove instances from rotation, potentially worsening the outage. Alerting on the unhealthy instance count would not measure the actual error rate experienced by users and could trigger false positives during maintenance. It also does not provide a 5% threshold on errors, and maintenance windows would not suppress alerts based on instance health.

  • ✓

    Create a log-based metric that counts HTTP 5xx responses from the load balancer logs, then create an alerting policy on that metric with a threshold of 5% error rate and configure a maintenance window for planned downtime.

    Why this is correct

    Log-based metrics derive values from log entries, and the load balancer logs include status details for each request. By counting 5xx responses and dividing by total requests, you can compute an error rate. Alerting policies can use such metrics with threshold conditions, and maintenance windows suppress alerts during planned downtime. This directly addresses the requirement to monitor application-level errors rather than infrastructure health, and respects maintenance periods.

About these practice questions

This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.