Courseiva

Google PCA Practice Question: Managing Implementation and Ensuring Solution and Operations Reliability

You are responsible for operations reliability of a production service running on Google Cloud. The service is deployed on GKE and exposes an external HTTPS endpoint through an external Application Load Balancer. You need to implement monitoring that detects when the service is unhealthy from the user's perspective and alerts the on-call team. (Choose two.)

⚠ Common exam trap

The trap here is choosing internal resource metrics, such as OOMKilled logs or span counts, as the primary health signal instead of external probes and edge-level error and latency metrics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set up an alerting policy on the load balancer's 5xx error rate and on backend latency, with thresholds tied to the service level objective.

Detecting user-facing unhealthiness requires probing the service as users reach it and monitoring the error and latency signals that reflect their experience. An uptime check from multiple locations validates the public endpoint and response content, while alerting on load balancer 5xx rates and backend latency ties detection to SLO thresholds. Together they catch both total outages and degraded performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create an alerting policy on the Application Load Balancer's request count metric to fire when traffic drops below a static threshold.

    Why it's wrong here

    Request count reflects traffic volume, not health. A drop in traffic could be caused by low demand, a client-side issue, or a load balancer misconfiguration, and a healthy service can receive little traffic. Alerting on traffic volume produces false positives and misses failures where requests still arrive but return errors.

  • ✗

    Enable Cloud Trace on the GKE workloads and alert when the number of spans per minute exceeds a fixed value.

    Why it's wrong here

    Cloud Trace collects distributed traces for latency analysis, but span volume is not a health indicator. A high span count may simply reflect increased traffic, and a failing service may emit fewer spans. Alerting on span volume does not reliably indicate user-facing failures and would generate noise rather than actionable alerts.

  • ✓

    Set up an alerting policy on the load balancer's 5xx error rate and on backend latency, with thresholds tied to the service level objective.

    Why this is correct

    Load balancer 5xx rates and backend latency metrics capture server-side failures and performance degradation as seen at the edge. Alerting on these signals against SLO-derived thresholds detects when users experience errors or slow responses. Combined with an uptime check, this provides both external probing and internal telemetry for reliable detection.

  • ✗

    Configure a log-based alert on GKE node system logs for the keyword 'OOMKilled'.

    Why it's wrong here

    Log-based alerts on OOMKilled events can reveal memory pressure in containers, but they do not measure user-visible availability. A node may log OOMKilled events while the service still responds correctly, or the service may be down without any OOMKilled event. This signal is useful for diagnostics but insufficient as the primary detection of user-facing unhealthiness.

  • ✓

    Create an uptime check in Cloud Monitoring that targets the external HTTPS URL and verifies the expected response code and content.

    Why this is correct

    An uptime check probes the service from multiple global locations over the public internet, which reflects the user's perspective. By validating the response code and expected content, it detects failures in the full path, including the load balancer, TLS, and backend. Alerting policies attached to the uptime check notify the on-call team when the service is unreachable or returns unexpected content.

About these practice questions

This PCA question is part of Courseiva's 807-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PCA practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PCA exam.