Courseiva

Cloud Digital Leader Scaling with Google Cloud operations Practice Question

An e-commerce platform uses Compute Engine instances in a managed instance group behind a Cloud Load Balancer. During a flash sale, the load balancer reports increased error rates. The operations team suspects the instances are overwhelmed. Which two steps should they take to troubleshoot the issue? (Choose TWO.)

⚠ Common exam trap

It's easy for candidates to confuse switching load balancer types (Option A) with a performance fix, when in fact the issue is backend capacity, not frontend protocol handling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check the CPU utilization of the instance group in Cloud Monitoring.

Option D is correct because checking CPU utilization of the instance group in Cloud Monitoring directly tests the hypothesis that the instances are overwhelmed; sustained high CPU across the MIG is the key signal of saturation and helps confirm whether the error rate spike is caused by insufficient compute capacity. Option E is correct because reviewing the load balancer logs in Cloud Logging exposes the actual HTTP error codes and messages (for example 502/503 responses or backend timeouts), which distinguishes backend overload from other causes such as misconfigured health checks or client errors. Option A is not appropriate because switching to a Network Load Balancer is a design change, not a troubleshooting step, and it would not diagnose why the current load balancer is reporting errors. Option B is not appropriate because blindly increasing the size of the instance group without investigation may mask the symptom, waste resources, and leave the root cause unidentified. Option C is not appropriate because enabling HTTP health checks is a configuration change rather than a diagnostic action, and health checks alone would not reveal the cause of the elevated error rates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch to a Network Load Balancer for higher throughput.

    Why it's wrong here

    Switching from an HTTP(S) load balancer to a Network Load Balancer (L4) alters the architectural layer and disables all L7 features such as path-based routing, TLS termination, and HTTP health checks. This change does not address random 5xx errors or high CPU utilization, because throughput is rarely the limiting factor when individual backend instances are already saturated. The appropriate first step is to inspect the existing load balancer's logs and backend metrics before considering any topology change.

  • ✗

    Increase the size of the instance group without investigation.

    Why it's wrong here

    Increasing the target size of the managed instance group prior to investigating the symptom risks adding unnecessary infrastructure that does not fix the root cause. For instance, if the error stems from an application bug, a database bottleneck, or a memory leak, additional Compute Engine vCPUs have no effect on those failure modes. Cloud Monitoring should first be used to validate whether CPU, memory, or network utilization is the true constraint that autoscaling would address.

  • ✗

    Enable HTTP health checks on the load balancer.

    Why it's wrong here

    Enabling HTTP health checks on the load balancer serves a preventive function by allowing it to stop routing traffic to unresponsive backend instances, but it is not a diagnostic tool for an ongoing incident. Even if health checks were misconfigured, simply toggling them introduces routing changes that could momentarily increase error rates and does not explain why instances became unhealthy. The immediate priority is to read the load balancer's logs and Cloud Monitoring metrics to identify the actual error source.

  • ✓

    Check the CPU utilization of the instance group in Cloud Monitoring.

    Why this is correct

    Inspect the Cloud Monitoring metrics for the managed instance group, specifically the CPU utilization time series (compute.googleapis.com/instance/cpu/utilization). Sustained high CPU percentages indicate that the existing instances are computing at their capacity, which directly supports the hypothesis that backend overload is producing errors or latency. This observation helps determine whether autoscaling or code optimization is the correct next step, rather than blindly changing infrastructure.

  • ✓

    Review the load balancer logs in Cloud Logging for error messages.

    Why this is correct

    Query Cloud Logging for HTTP access logs emitted by the external HTTP(S) load balancer, filtering on 5xx response codes and backend response time. These logs reveal whether the origin of the failures is the load balancer itself or the Compute Engine backends, and they pinpoint which request paths exhibit elevated latency or errors. Examining this structured data is the most direct way to identify the exact HTTP error pattern driving the user-facing issue.

Visual reference

192.168.1.0 /24 256 addresses (254 usable) 192.168.1.0 /25 Subnet A 128 addr (126 usable) 192.168.1.128 /25 Subnet B 128 addr (126 usable) Borrowing 1 bit from host portion creates 2 subnets (/25)

About these practice questions

One of 848 original GCDL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.