Courseiva

CCNA Ensure solution and operations reliability Questions

58 questions · Ensure solution and operations reliability · All types, answers revealed

1
MCQmedium

Your company's global e-commerce platform uses a managed instance group (MIG) in us-central1 and a Cloud Load Balancer. Traffic has grown, and you want to improve availability by distributing load across multiple regions. What should you do?

A.Increase the machine type of the existing instances to handle more traffic.
B.Enable Cloud CDN to cache content closer to users.
C.Create MIGs in additional regions and add them as backends to the existing global load balancer.
D.Change the load balancer to global and configure a single backend.
AnswerC

Adding regional managed instance groups as backends to the existing global load balancer satisfies the multi-region availability requirement. A global external Application Load Balancer routes users to the closest healthy backend and performs cross-region failover, so capacity in additional regions absorbs traffic when one region degrades.

Why this answer

A global external HTTP(S) load balancer can have backends in multiple regions. By creating managed instance groups (MIGs) in additional regions and adding them as backends to the existing global load balancer, you distribute traffic across regions, improving availability and reducing latency for users worldwide. This approach leverages the load balancer's anycast IP and cross-region load balancing capabilities.

Exam trap

The trap here is that candidates confuse Cloud CDN (which caches content) with multi-region backend distribution, or think that simply making the load balancer 'global' with a single backend achieves regional redundancy, when in fact you must add backends in multiple regions to distribute load and improve availability.

How to eliminate wrong answers

Option A is wrong because increasing the machine type of existing instances only scales vertically within a single region, which does not address multi-region availability or distribute load geographically. Option B is wrong because Cloud CDN caches static content at edge locations but does not distribute compute load across regions; it reduces latency for cached content but does not improve availability for dynamic requests or handle regional failures. Option D is wrong because changing the load balancer to global and configuring a single backend (a single MIG) still limits compute resources to one region, failing to provide multi-region distribution or fault isolation.

2
Multi-Selecteasy

A company wants to monitor the health of their Cloud Run services. Which THREE metrics should they use to define a comprehensive health SLI? (Choose 3)

Select 3 answers
A.Latency (e.g., p99 response time)
B.CPU utilization
C.Request count
D.Instance count
E.Error rate (percentage of 5xx responses)
AnswersA, C, E

Latency is a key performance SLI for user experience.

Why this answer

Latency (p99 response time) is a critical metric for Cloud Run because it measures the end-to-end request processing time, directly reflecting user experience. In a serverless environment, high latency can indicate cold starts, insufficient concurrency, or downstream service bottlenecks, making it essential for a comprehensive health SLI.

Exam trap

Google Cloud often tests the misconception that infrastructure-level metrics like CPU or instance count are valid health SLIs for serverless services, when in fact user-facing metrics (latency, errors, request count) are the correct choices for a comprehensive health SLI.

3
MCQeasy

A company uses Cloud SQL for PostgreSQL. They want to minimize downtime during maintenance. Which feature should they enable?

A.Read replicas.
B.High availability with a standby in another zone.
C.Point-in-time recovery.
D.Automated backups.
AnswerB

High availability provisions a standby replica in a different zone; during maintenance or a zone failure, Cloud SQL fails over automatically, minimising downtime. Read replicas and automated backups do not provide this automatic cross-zone failover, so they cannot satisfy the stem's downtime constraint.

Why this answer

High availability (HA) with a standby in another zone ensures that Cloud SQL for PostgreSQL automatically fails over to a standby instance in a different zone if the primary zone experiences an outage. This minimizes downtime during maintenance because Cloud SQL performs a controlled failover to the standby, typically completing within a few seconds, rather than requiring a full instance restart or rebuild.

Exam trap

The trap here is that candidates often confuse read replicas with high availability, assuming read replicas can automatically take over for the primary, but read replicas require manual promotion and do not provide automatic failover, making HA with a standby the correct choice for minimizing downtime during maintenance.

How to eliminate wrong answers

Option A is wrong because read replicas are designed for offloading read traffic and do not provide automatic failover for the primary instance; they require manual promotion, which introduces downtime. Option C is wrong because point-in-time recovery (PITR) is used for restoring data to a specific timestamp after data corruption or accidental deletion, not for reducing downtime during planned maintenance. Option D is wrong because automated backups protect against data loss by creating periodic backups, but they do not provide a standby instance for failover, so maintenance still requires downtime to restart the primary instance.

4
MCQeasy

Your company runs a critical application on Google Kubernetes Engine (GKE) with 5 nodes. The application experiences intermittent high latency every Friday afternoon. The team has ruled out infrastructure issues and suspects the application logic. You need to instrument the application to identify the root cause. Which approach should you take?

A.Use Cloud Monitoring to create custom metrics for application performance and investigate recent code changes.
B.Increase the number of nodes in the GKE cluster to handle the load.
C.Enable Cloud Logging and analyze logs for error messages during the latency periods.
D.Configure GKE usage metering to track resource consumption by namespace.
AnswerA

Cloud Monitoring custom metrics expose application-level performance signals, letting the team correlate Friday latency spikes with recent code changes rather than infrastructure. This instruments application logic directly, satisfying the need to pinpoint the root cause after infrastructure was ruled out.

Why this answer

The team has already ruled out infrastructure issues and suspects application logic. Creating custom metrics in Cloud Monitoring allows you to instrument the application with key performance indicators (e.g., request latency, error rates) and correlate them with recent code changes to pinpoint the root cause of intermittent high latency. This approach directly addresses the need to monitor application-level behavior rather than infrastructure metrics.

Exam trap

The trap here is that candidates often confuse operational logging (Option C) with performance monitoring, failing to recognize that intermittent latency without errors requires custom metrics to measure application-specific performance indicators.

How to eliminate wrong answers

Option B is wrong because increasing the number of nodes addresses infrastructure capacity, which has already been ruled out as the cause; it does not help identify application logic issues. Option C is wrong because while Cloud Logging can capture error messages, the problem is intermittent high latency without necessarily generating errors; analyzing logs alone may miss performance bottlenecks that require custom metrics. Option D is wrong because GKE usage metering tracks resource consumption by namespace for cost allocation, not application performance or latency issues.

5
MCQhard

A financial services firm runs a latency-sensitive trading API on Google Kubernetes Engine (GKE). During peak market hours, the API occasionally returns errors because pods are evicted when nodes run out of memory. The team wants the workload to be protected from node-level resource pressure and to receive a graceful termination window when the node must be drained. Which configuration should they apply to the Deployment?

A.Configure the pods with Guaranteed QoS by setting requests equal to limits, and set a high priorityClassName so they are evicted last.
B.Set a PodDisruptionBudget with minAvailable equal to the replica count and configure a longer terminationGracePeriodSeconds.
C.Set memory requests equal to limits to achieve Guaranteed QoS, assign a high-priority class, and increase terminationGracePeriodSeconds to allow graceful shutdown.
D.Define resource requests and limits for CPU and memory, and set the pod's priorityClassName to a high-priority class.
AnswerC

Guaranteed QoS makes the pods least likely to be evicted under node memory pressure because they are treated as the highest-priority QoS class. A high-priority class further ensures they are evicted after lower-priority workloads. The longer terminationGracePeriodSeconds gives the container time to finish in-flight trades before SIGKILL, addressing both the eviction protection and graceful-termination requirements in the scenario.

Why this answer

Guaranteed QoS, achieved by setting memory and CPU requests equal to limits, makes pods the last to be evicted when a node is under memory pressure. Pairing that with a high-priority class and a longer terminationGracePeriodSeconds protects the trading API from involuntary eviction and gives it time to drain in-flight requests gracefully, matching both stated requirements.

Exam trap

The trap here is treating a PodDisruptionBudget as protection against node memory-pressure evictions, when it only governs voluntary disruptions such as drains and upgrades.

6
Multi-Selectmedium

A company has deployed a critical application on Google Kubernetes Engine (GKE) with a Regional cluster (us-central1). The application uses a Cloud SQL for PostgreSQL database with a cross-region replica for disaster recovery. The SRE team needs to ensure that the application can survive a regional outage with minimal data loss. Which TWO actions should the team take to improve the reliability of the solution?

Select 2 answers
A.Configure the application to automatically promote the Cloud SQL cross-region replica to a primary instance when the primary region is unavailable.
B.Configure Cloud SQL cross-region replication to be synchronous to ensure zero data loss during failover.
C.Configure an external HTTP(S) load balancer with a backend service pointing to both the primary and secondary GKE clusters, and use a DNS failover policy to route traffic to the secondary region if the primary region becomes unhealthy.
D.Deploy a secondary GKE cluster in the same region as the primary to provide a hot standby that can take over immediately.
E.Use a TCP/UDP load balancer to route traffic to both regions based on latency.
AnswersA, C

Correct. Promoting a Cloud SQL cross-region replica to a primary instance is the standard DR procedure. Automation via Cloud Functions or Cloud Run reduces RTO. However, cross-region replication is asynchronous, so some data loss is possible.

Why this answer

To survive a regional outage with minimal data loss, two key actions are needed: (1) Automate promotion of the Cloud SQL cross-region replica to primary when the primary region fails (Option A) – this is the standard Cloud SQL DR procedure; automation via Cloud Functions/Cloud Run reduces RTO. Note that cross-region replication is asynchronous, so some data loss (replication lag) is possible. (2) Deploy a multi-region GKE cluster pair (primary and secondary) and use an external HTTP(S) load balancer with a backend service pointing to both clusters, combined with a DNS failover policy (Option C) – this allows the load balancer to detect primary region health and route traffic to the secondary region if needed. Options B is wrong because cross-region replication cannot be synchronous.

Option D is wrong because a secondary cluster in the same region does not help during a regional outage. Option E is wrong because a TCP/UDP load balancer with latency-based routing does not provide failover based on region health.

Exam trap

The trap here is that candidates often assume synchronous replication is possible across regions for zero data loss, but in practice, cross-region replication is always asynchronous due to the speed of light and network latency constraints.

7
MCQmedium

A company monitors their application with Cloud Monitoring. They set up an alerting policy to notify the on-call team when the 99th percentile latency exceeds 500 ms for 5 minutes. However, they receive false positive alerts due to short bursts. How should they refine the policy?

A.Set up alerting on each data point individually.
B.Decrease the threshold to 400 ms.
C.Change the metric to average latency instead of 99th percentile.
D.Increase the evaluation window to 10 minutes.
AnswerD

Extending the evaluation window to 10 minutes requires latency to breach 500 ms across a longer sustained period, filtering out brief spikes that triggered false positives. This directly addresses the stem's short-burst problem while retaining detection of genuine sustained degradation.

Why this answer

Increasing the evaluation window to 10 minutes smooths out short bursts of high latency, ensuring the alert triggers only when the 99th percentile latency exceeds 500 ms for a sustained period. Cloud Monitoring evaluates metrics over the specified window, so a longer window reduces false positives from transient spikes while still detecting genuine degradation.

Exam trap

Google Cloud often tests the misconception that lowering thresholds or changing percentiles reduces false positives, when in reality the evaluation window duration is the key lever for filtering out short-lived bursts without sacrificing sensitivity to sustained issues.

How to eliminate wrong answers

Option A is wrong because setting up alerting on each data point individually would make the policy hypersensitive to every single spike, increasing false positives rather than reducing them. Option B is wrong because decreasing the threshold to 400 ms would cause the alert to fire even more frequently, including during normal operation, exacerbating the false positive problem. Option C is wrong because changing the metric to average latency masks tail latency issues; the 99th percentile is specifically used to catch outliers, and averaging would hide the very bursts they want to monitor, potentially missing real problems.

8
MCQeasy

A company uses Cloud Storage for backup data. They want to protect against accidental deletion. Which option is best?

A.Enable object versioning.
B.Use a lifecycle policy.
C.Set a retention policy.
D.Use object holds.
AnswerA

Preserves noncurrent versions for recovery.

Why this answer

Object versioning in Cloud Storage preserves every version of an object, including overwrites and deletions. When versioning is enabled, a delete operation creates a delete marker instead of permanently removing the object, allowing easy recovery. This directly protects against accidental deletion by retaining all previous object versions.

Exam trap

Google Cloud often tests the distinction between versioning (which allows recovery from accidental deletion) and retention policies (which prevent deletion but do not provide recovery after the fact), leading candidates to confuse compliance protection with accidental deletion protection.

How to eliminate wrong answers

Option B is wrong because lifecycle policies automate transitions or deletions based on age or conditions, but they do not prevent accidental deletion; they can actually cause deletion if misconfigured. Option C is wrong because retention policies (e.g., Bucket Lock) prevent object modification or deletion for a fixed period, but they are designed for compliance and data retention, not for recovering from accidental deletion after the fact. Option D is a duplicate of the correct answer and is not a separate option; the question lists two identical 'Enable object versioning' entries, but only one is correct.

9
MCQeasy

A company runs a global e-commerce site on GKE. They want to ensure disaster recovery with multi-region deployment. What is the best practice for configuring GKE clusters?

A.Deploy separate regional clusters in two or more regions.
B.Use a single zonal cluster with node auto-repair.
C.Deploy a single cluster with multi-master setup.
D.Use a single regional cluster with multiple zones.
AnswerA

Separate regional clusters place control planes and nodes in distinct regions, so a single region's failure leaves the other serving traffic. This satisfies the multi-region disaster recovery constraint, unlike zonal clusters or single-region node pools.

Why this answer

For disaster recovery with a multi-region deployment, the best practice is to deploy separate regional clusters in two or more regions. This ensures that if an entire region fails, traffic can be redirected to the other region's cluster, providing true geographic redundancy. A single cluster, whether zonal or regional, cannot survive a regional outage because it is bound to a single control plane location.

Exam trap

Google Cloud often tests the misconception that a regional cluster with multiple zones is sufficient for disaster recovery, but the trap here is that a regional cluster is still confined to a single region and cannot survive a full regional outage.

How to eliminate wrong answers

Option B is wrong because a single zonal cluster with node auto-repair only protects against node-level failures within that single zone, not against a full zone or regional outage, and thus does not meet multi-region disaster recovery requirements. Option C is wrong because GKE does not support a multi-master setup; each cluster has a single control plane, and multi-master is not a valid configuration for GKE. Option D is wrong because a single regional cluster with multiple zones provides high availability within a single region but cannot survive a regional failure, as the control plane is still regional and would be unavailable if the entire region goes down.

10
MCQeasy

A logistics company runs a Cloud Run service that processes shipment events. They want to be notified and to trigger an automated rollback when the error rate of a new revision exceeds a threshold shortly after deployment. Which Google Cloud feature should they use?

A.Use Error Reporting to group exceptions and configure a Pub/Sub notification that emails the on-call engineer to perform a rollback.
B.Enable Cloud Run's built-in automatic rollback by setting a maximum error rate in the service YAML.
C.Configure Cloud Run gradual rollout with canary traffic splitting and manually monitor the revision before shifting all traffic.
D.Create a Cloud Monitoring alerting policy on the Cloud Run error rate and use Cloud Deploy with a deployment verification and automated rollback.
AnswerD

Cloud Deploy supports deployment strategies with verification, where a Cloud Monitoring alert or custom job evaluates the new revision and, on failure, automatically rolls back to the previous stable release. This provides both the notification via the alerting policy and the automated rollback the team wants, without manual intervention, matching the stated requirement precisely.

Why this answer

Cloud Deploy deployment verification evaluates a new Cloud Run revision against defined criteria, such as a Cloud Monitoring alert on error rate, and automatically rolls back to the prior stable release when verification fails. Pairing it with a Cloud Monitoring alerting policy delivers both the notification and the automated rollback the logistics team requires.

Exam trap

The trap here is assuming Cloud Run has a native error-rate-based automatic rollback, when automated rollback is orchestrated through Cloud Deploy verification instead.

11
MCQeasy

Your company runs a stateless web application on Compute Engine. You want to ensure that if a zone fails, the application continues to serve traffic with minimal manual intervention. What should you do?

A.Schedule regular snapshots of each instance's persistent disk to a regional bucket.
B.Create a regional managed instance group with an autoscaling policy and use a global Cloud Load Balancer.
C.Use a global Cloud Load Balancer and enable Cloud CDN.
D.Create an instance template and manually deploy instances in another zone.
AnswerB

A regional managed instance group spreads instances across multiple zones, so a zone failure leaves capacity elsewhere, while the global Cloud Load Balancer directs traffic to healthy backends. Autoscaling maintains capacity, meeting the minimal-manual-intervention requirement for the stateless application.

Why this answer

A regional managed instance group (MIG) distributes instances across multiple zones within a region, ensuring that if one zone fails, the remaining zones continue serving traffic. Combined with a global Cloud Load Balancer, traffic is automatically routed to healthy instances in any zone, providing high availability with minimal manual intervention. Autoscaling further ensures that new instances are created to handle load, even if a zone becomes unavailable.

Exam trap

Google Cloud often tests the distinction between data backup (snapshots) and compute redundancy (MIGs), leading candidates to choose backup solutions when the question asks for continuous traffic serving during a zone failure.

How to eliminate wrong answers

Option A is wrong because scheduling snapshots to a regional bucket provides data backup and disaster recovery for persistent disks, but does not automatically redirect traffic or maintain application availability during a zone failure; it requires manual restoration and reconfiguration. Option C is wrong because enabling Cloud CDN caches static content at edge locations, which improves performance and reduces load on origin servers, but does not provide zone-level redundancy or automatic failover for the compute instances themselves. Option D is wrong because manually deploying instances in another zone is a manual, slow process that does not provide automated failover or load balancing; it also lacks autoscaling and health checking, leading to potential downtime and increased operational overhead.

12
Multi-Selecteasy

A company uses Cloud Storage to store user-uploaded content. They want to ensure that the data is highly durable and protected against accidental deletion. Which two features should they enable? (Choose two.)

Select 2 answers
A.Requester pays.
B.Lifecycle management.
C.Object versioning.
D.Bucket retention policy.
E.Uniform bucket-level access.
AnswersC, D

Object versioning preserves every revision of an object, so overwrites and deletions create non-current versions rather than destroying data. This directly satisfies the accidental-deletion protection requirement, allowing recovery of prior content. Combined with retention or lifecycle rules, it guards against both user error and malicious removal.

Why this answer

Object versioning (C) is correct because it keeps prior versions of objects when they are overwritten or deleted, so an accidentally deleted or replaced object can be recovered rather than permanently lost. Bucket retention policy (D) is correct because it enforces a retention period during which objects cannot be deleted or overwritten, directly protecting data against accidental or premature deletion. Together these features address the durability and deletion-protection requirement for user-uploaded content in Cloud Storage.

Requester pays (A) only shifts data-access and egress costs to the requester and provides no deletion protection. Lifecycle management (B) automates actions such as deleting or transitioning objects and can actually cause deletion, so it does not protect against accidental loss. Uniform bucket-level access (E) simplifies permission management by applying IAM uniformly at the bucket level, but it does not prevent accidental object deletion.

13
MCQhard

You are deploying a new version of a microservice to Google Kubernetes Engine (GKE). The service must remain available during the rollout, and you need to minimize the risk of exposing bugs to all users at once. You want to gradually shift traffic to the new version while monitoring key metrics. Which strategy should you use?

A.Implement a canary deployment using Istio or Anthos Service Mesh to route a small percentage of traffic to the new version.
B.Perform a rolling update by changing the Deployment's pod template.
C.Use a blue/green deployment by creating a new Deployment and switching the Service selector.
D.Create a new GKE cluster and deploy the new version there, then update DNS to point to the new cluster.
AnswerA

A canary deployment with a service mesh like Istio allows you to route a small percentage of traffic to the new version, monitor metrics, and gradually increase traffic or roll back if issues arise. This minimizes risk and keeps the service available. It provides the fine-grained control needed for safe rollouts.

Why this answer

A canary deployment using a service mesh allows you to route a small portion of traffic to the new version, monitor its behavior, and then gradually increase traffic or roll back if problems occur. This minimizes risk and maintains availability. Other strategies either switch all traffic at once or lack the granular control needed for gradual, metric-based rollouts.

Exam trap

The trap here is confusing a rolling update with a canary deployment; rolling updates replace pods but do not control traffic percentages or support metric-based analysis.

14
Drag & Dropmedium

Drag and drop the steps to set up a Cloud VPN tunnel between Google Cloud and an on-premises network into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Cloud Router is used for dynamic routing. The tunnel requires the on-premises public IP and pre-shared key.

15
MCQeasy

A company wants to monitor their Cloud Run services for errors and latency. Which Google Cloud product should they use?

A.Cloud Trace
B.Cloud Monitoring
C.Cloud Logging
D.Error Reporting
AnswerB

Cloud Monitoring natively ingests Cloud Run request metrics, error rates and latency percentiles, letting the company alert on them without custom instrumentation. It satisfies the stated requirement to monitor errors and latency across their services, unlike logging-only or CI/CD tools.

Why this answer

Cloud Monitoring (formerly Stackdriver Monitoring) provides comprehensive observability for Cloud Run services, including built-in dashboards for request latency, error rates, and resource utilization. It collects metrics like request count, request latencies, and container instance counts, and allows you to set alerting policies based on these metrics. While Cloud Trace can help with latency analysis and Cloud Logging captures logs, Cloud Monitoring is the primary product for monitoring both errors and latency in a unified view.

Exam trap

The trap here is that candidates often confuse Cloud Trace (for latency) or Error Reporting (for errors) as standalone solutions, but the question asks for a single product that monitors both errors and latency, which is Cloud Monitoring's role as the central metrics and alerting platform.

How to eliminate wrong answers

Option A is wrong because Cloud Trace is a distributed tracing tool focused on analyzing latency across service requests, but it does not provide a unified dashboard for error rates or resource metrics for Cloud Run. Option C is wrong because Cloud Logging is for storing, searching, and analyzing log data, not for monitoring metrics like latency percentiles or error counts in real-time dashboards. Option D is wrong because Error Reporting aggregates and analyzes application errors from logs, but it does not monitor latency or provide a holistic view of service health.

16
Multi-Selecthard

Your service has a 99.99% uptime SLO (monthly error budget ~ 4 minutes). Which TWO monitoring practices best support this SLO? (Choose 2)

Select 2 answers
A.Monitor CPU utilization and alert when average exceeds 80%.
B.Use a combination of availability (e.g., HTTP 200 rate) and latency (e.g., p99) as SLIs.
C.Use only synthetic monitoring from multiple locations.
D.Alert on every 5xx error immediately.
E.Track error budget consumption and alert when burn rate exceeds a threshold.
AnswersB, E

Availability alone misses degraded-but-successful responses, which still breach user expectations on a 99.99% target. Pairing HTTP 200 rate with p99 latency captures both failure and slowness, so the SLI reflects the actual user experience the SLO promises.

Why this answer

Option B is correct because a 99.99% uptime SLO is best measured with SLIs that reflect user-perceived health, so combining an availability SLI (the proportion of successful HTTP 200 responses) with a latency SLI (such as p99 request duration) captures both whether requests succeed and whether they are served acceptably fast. Option E is correct because with a monthly error budget of roughly 4 minutes, tracking error budget consumption and alerting on a burn rate threshold (for example, a fast-burn alert at 14.4x over 1 hour or a slow-burn alert at 6x over 6 hours) detects when the budget is being exhausted too quickly and enables timely action. Option A is not correct because CPU utilization is a resource metric, not a direct SLI for an uptime SLO, and an 80% average threshold does not reliably indicate user-visible failures.

Option C is not correct because relying only on synthetic monitoring from multiple locations omits real user traffic and can miss failures that affect actual customers. Option D is not correct because alerting on every 5xx error immediately is too noisy and does not account for error budget policy or burn rate.

Exam trap

PCA often tests the difference between resource metrics (CPU, memory) and user-centric SLIs — candidates pick CPU alerts because they are familiar, missing that SLOs must be measured with user-facing indicators and error budget burn.

17
Multi-Selecthard

A company runs a stateful application on GKE using StatefulSets. Which THREE practices improve reliability?

Select 3 answers
A.Use headless services.
B.Use horizontal autoscaling based on disk usage.
C.Use volume snapshots for backup.
D.Use pod disruption budgets.
E.Use persistent volumes with reclaim policy Delete.
AnswersA, C, D

Provides stable network identities for stateful workloads.

Why this answer

A headless service (clusterIP: None) allows direct pod-to-pod communication without load balancing, which is essential for stateful applications like databases that require stable network identities. Each pod in a StatefulSet gets a unique DNS name (e.g., pod-0.service.namespace.svc.cluster.local), enabling reliable discovery and ordering for replication, leader election, and failover. This ensures that clients always reach the correct pod instance, improving overall reliability.

Exam trap

Google Cloud often tests the misconception that horizontal autoscaling can be based on any arbitrary metric like disk usage, but the HPA only supports CPU, memory, and custom/external metrics that must be exposed through the Metrics Server or a custom metrics adapter.

18
MCQhard

You are running a Kubernetes cluster in GKE with the default node pool configuration shown in the exhibit. Your application requires high disk I/O performance. You notice that the application is experiencing high latency for disk operations. What is the most likely cause?

A.Node auto-repair is causing disk contention.
B.The default node pool uses pd-standard disks, which have low IOPS.
C.The OAuth scopes restrict disk access, causing high latency.
D.The machine type n1-standard-2 does not have enough CPU.
AnswerB

GKE's default node pool provisions pd-standard persistent disks, which cap IOPS far below pd-ssd or pd-balanced. Since the stem specifies high disk I/O demand and observed latency, the storage class backing the nodes is the bottleneck. Upgrading the node pool to pd-ssd resolves the constraint.

Why this answer

The default node pool in GKE uses pd-standard (standard persistent disk) which provides lower IOPS compared to pd-ssd. For applications requiring high disk I/O performance, pd-standard disks become a bottleneck, causing high latency. Upgrading to pd-ssd or using local SSDs would resolve this issue.

Exam trap

Google Cloud often tests the distinction between storage performance (disk type) and other operational features (auto-repair, scopes, machine type), leading candidates to confuse node health mechanisms or permission settings with actual I/O performance bottlenecks.

How to eliminate wrong answers

Option A is wrong because node auto-repair is a GKE feature that automatically repairs unhealthy nodes (e.g., if the node fails health checks), but it does not cause disk contention; it operates at the node level, not by interfering with disk I/O. Option C is wrong because OAuth scopes control API access permissions (e.g., read/write to Cloud Storage), not the performance characteristics of persistent disk operations; disk I/O latency is a storage performance issue, not an authorization issue. Option D is wrong because n1-standard-2 (2 vCPUs, 7.5 GB memory) is a general-purpose machine type that can handle moderate workloads; insufficient CPU would manifest as high CPU utilization or scheduling delays, not specifically high disk I/O latency.

19
MCQhard

Your company runs a critical multi-tier application: a global HTTP(S) load balancer, multiple regional managed instance groups (MIGs) for the web tier, and Cloud Spanner for the data tier. You need to design for zone-level and region-level failures. What architecture ensures the highest availability?

A.Use a global HTTP(S) load balancer with a single global MIG and a multi-region Cloud Spanner instance.
B.Use a global HTTP(S) load balancer with a single zonal MIG and Cloud Spanner single-region.
C.Use a global HTTP(S) load balancer with regional MIGs in multiple regions, each spanning zones, and a multi-region Cloud Spanner instance.
D.Use a regional HTTP(S) load balancer with a regional MIG and Cloud SQL with cross-region replication.
AnswerC

Regional MIGs spanning zones absorb zone-level failures, while distributing them across multiple regions satisfies region-level resilience. The global HTTP(S) load balancer anycasts traffic to the nearest healthy backend, and the multi-region Cloud Spanner instance provides synchronous cross-region replication with automatic failover, meeting both failure scopes in the stem.

Why this answer

It combines a global HTTP(S) load balancer (which can route traffic to healthy backends across regions), regional MIGs that span multiple zones within each region (providing zone-level redundancy), and a multi-region Cloud Spanner instance (which provides synchronous replication across regions for strong consistency and automatic failover). This architecture ensures that if an entire zone or region fails, traffic is automatically redirected to healthy backends in other zones/regions, and Spanner continues to serve reads and writes without manual intervention.

Exam trap

Google Cloud often tests the distinction between 'regional' and 'global' load balancers, and the trap here is that candidates might choose a regional load balancer (Option D) thinking it is sufficient, but it cannot route traffic across regions, making it unsuitable for region-level failure recovery.

How to eliminate wrong answers

Option A is wrong because a single global MIG (even if multi-zonal) is still deployed within a single region; if that entire region fails, the application becomes unavailable. Option B is wrong because a single zonal MIG cannot survive even a zone failure, and a single-region Cloud Spanner instance cannot survive a regional failure. Option D is wrong because a regional HTTP(S) load balancer cannot distribute traffic across multiple regions, and Cloud SQL with cross-region replication does not provide the same strong consistency and automatic failover as multi-region Spanner; also, Cloud SQL cross-region replication is asynchronous and may lose data during a failover.

20
MCQmedium

A retail company uses a Cloud SQL for MySQL instance with a single zone. The database is critical for order processing, and the company wants to minimize downtime if the zone hosting the instance fails. They also want to ensure that the application can continue to write data during a zonal failure without manual intervention. What should they do?

A.Migrate the database to a Compute Engine instance with a regional persistent disk.
B.Create a read replica in another zone and promote it if the primary fails.
C.Enable high availability by configuring the instance as regional.
D.Take regular automated backups and restore to a new instance in another zone.
AnswerC

Configuring a Cloud SQL instance as regional creates a standby replica in a different zone. If the primary zone fails, Cloud SQL automatically fails over to the standby, allowing writes to continue with minimal downtime. This matches the requirement for automatic failover and continued write availability during a zonal failure.

Why this answer

Configuring Cloud SQL as regional provides a synchronous standby in another zone and automatic failover. This ensures that writes can continue with minimal downtime during a zonal failure, satisfying the business requirement. Other options either require manual intervention, result in data loss, or do not provide automatic failover.

Exam trap

The trap here is assuming that a read replica can provide high availability, but read replicas do not offer automatic failover and are intended for read scaling.

21
MCQhard

A healthcare company runs a patient portal on Cloud Run services in the us-central1 region. The compliance team requires that the portal remain readable during a regional outage and that failover to a secondary region occur without changing the public hostname. The portal's data is stored in Cloud SQL for PostgreSQL. Which design should the architect recommend?

A.Deploy the Cloud Run services in both regions and use Cloud DNS geolocation routing with a 300-second TTL to direct users to the healthy region.
B.Deploy the Cloud Run services in us-central1 only and front them with a regional external Application Load Balancer and a Cloud DNS failover routing policy pointing to a static IP.
C.Deploy the Cloud Run services in both regions behind a global external Application Load Balancer, and rely on Cloud SQL's automatic regional failover of the primary instance.
D.Deploy the Cloud Run services in us-central1 and us-east1 behind a global external Application Load Balancer with a serverless network endpoint group per region, and use a Cloud SQL cross-region replica promoted on failover.
AnswerD

A global external Application Load Balancer provides a single anycast hostname and can route to serverless network endpoint groups in multiple regions, so failover happens without DNS changes. A Cloud SQL cross-region replica keeps a warm copy of the data that can be promoted if the primary region fails, satisfying the readability requirement during a regional outage.

Why this answer

A global external Application Load Balancer with serverless network endpoint groups in two regions gives a stable anycast hostname that keeps working when one region fails, and Cloud Run's multi-region deployment provides compute in both places. For the data tier, a Cloud SQL cross-region replica must be promoted during regional failover, since Cloud SQL does not automatically fail over across regions.

Exam trap

The trap here is assuming Cloud SQL performs automatic cross-region failover like its within-region high availability mode, when cross-region recovery requires promoting a replica.

22
MCQeasy

You are the lead cloud architect for a startup that runs a web application on Google Kubernetes Engine (GKE) with a standard (zonal) cluster. The application is deployed with 3 replicas of a stateless frontend service. During a recent incident, a zone outage caused all GKE nodes to become unavailable, leading to application downtime of 45 minutes. You need to redesign the cluster to tolerate a single zone failure with no more than 5 minutes of downtime. Your budget allows for at most a 20% increase in compute costs. Which approach should you take?

A.Increase the number of replicas from 3 to 9 and keep the zonal cluster
B.Change the frontend deployment to use regional persistent disks
C.Deploy second GKE cluster in another region and use global load balancer for failover
D.Migrate the cluster to a regional GKE cluster with nodes in 3 zones and distribute replicas across zones
AnswerD

Correct: regional cluster survives zone failure.

Why this answer

D is correct because a regional GKE cluster distributes nodes across three zones, ensuring that if one zone fails, the remaining two zones continue serving traffic. By spreading the 3 replicas across zones (e.g., one per zone), the application tolerates a single zone outage with near-zero downtime, and the 20% cost increase covers the additional node pool overhead without exceeding the budget.

Exam trap

The trap here is that candidates confuse increasing replica count with achieving zone redundancy, failing to realize that replicas must be distributed across failure domains (zones) to survive a zone outage, and that regional persistent disks are irrelevant for stateless workloads.

How to eliminate wrong answers

Option A is wrong because increasing replicas to 9 in a zonal cluster does not provide zone redundancy; all nodes remain in a single zone, so a zone outage still takes down all replicas. Option B is wrong because regional persistent disks are used for stateful workloads (e.g., databases) and do not help with zone-level node failure for a stateless frontend; the frontend does not require persistent disks. Option C is wrong because deploying a second cluster in another region introduces cross-region latency and failover complexity, and the 5-minute downtime target cannot be met with DNS propagation or global load balancer failover; it also likely exceeds the 20% cost increase due to full cluster duplication.

23
MCQhard

The exhibit shows a managed instance group configuration. What is the primary purpose of the 'autoHealingPolicies' section?

A.Distribute incoming traffic evenly across the instances.
B.Automatically add more instances when CPU utilization exceeds 60%.
C.Automatically replace instances that are deemed unhealthy based on the health check.
D.Automatically update instances to a new instance template.
AnswerC

Autohealing policies continuously run the specified health check against each instance; any instance failing it is deleted and recreated by the managed instance group, restoring the desired capacity. This directly satisfies the stem's requirement to remediate unhealthy instances automatically, rather than merely scaling or distributing traffic.

Why this answer

The 'autoHealingPolicies' section in a managed instance group configuration is specifically designed to automatically replace instances that are deemed unhealthy based on a configured health check. When a health check probe (e.g., HTTP, TCP, or SSL) fails for a sustained period, the managed instance group terminates the unhealthy instance and creates a new one from the instance template, ensuring the desired number of healthy instances is maintained. This is distinct from autoscaling, which adjusts instance count based on load metrics.

Exam trap

Google Cloud often tests the distinction between 'autohealing' (health-based instance replacement) and 'autoscaling' (metric-based instance count adjustment), causing candidates to confuse the purpose of the 'autoHealingPolicies' section with scaling policies.

How to eliminate wrong answers

Option A is wrong because distributing incoming traffic evenly across instances is the function of a load balancer (e.g., HTTP(S) Load Balancer or Network Load Balancer) and its backend service, not the 'autoHealingPolicies' section of a managed instance group. Option B is wrong because automatically adding instances when CPU utilization exceeds 60% is a function of the 'autoscaling' policy (based on a CPU utilization metric), not the 'autoHealingPolicies' section, which only reacts to health check failures. Option D is wrong because automatically updating instances to a new instance template is achieved through a 'rolling update' or 'canary update' strategy (e.g., using the 'updatePolicy' section), not through 'autoHealingPolicies', which only replaces unhealthy instances with the current template.

24
Multi-Selectmedium

A company runs a stateful workload on Compute Engine with regional persistent disks (PD). They need to implement a disaster recovery (DR) plan with a Recovery Point Objective (RPO) of less than 1 hour and Recovery Time Objective (RTO) of less than 4 hours. Which THREE steps should they include in their DR plan? (Choose three.)

Select 3 answers
A.Take snapshots of the persistent disk every 30 minutes and copy them to a Cloud Storage bucket in another region
B.Create a snapshot schedule for the persistent disk every 4 hours
C.Create a custom machine image of the instance and store it in a Cloud Storage bucket in the DR region
D.Use regional persistent disks to automatically replicate data to a second zone
E.Test the failover procedure quarterly to validate RTO and RPO
AnswersA, C, E

Correct: meets RPO and protects against regional failure.

Why this answer

Taking snapshots every 30 minutes meets the RPO of less than 1 hour. By copying these snapshots to a Cloud Storage bucket in another region, you ensure data is available in a DR region for recovery, which is essential for cross-region disaster recovery.

Exam trap

The trap here is confusing zonal replication (regional PD) with cross-region disaster recovery; regional PDs only protect against zonal failures, not regional outages, so they cannot meet a cross-region DR requirement.

25
MCQhard

An organization uses Cloud Functions (2nd gen) for event-driven processing. They notice that some functions fail with 'memory limit exceeded' errors during peak load. The function processes messages from Pub/Sub and writes to Firestore. What should they do to improve reliability without sacrificing throughput?

A.Increase the maximum number of concurrent function instances.
B.Increase the memory allocated to the Cloud Function.
C.Enable Pub/Sub batching to reduce the number of function invocations.
D.Split the function into multiple smaller functions, each handling a subset of the data.
AnswerB

Memory limit exceeded errors mean the function's allocated memory is exhausted during peak concurrency. Raising the memory allocation gives each instance more headroom to process Pub/Sub messages and write to Firestore, restoring reliability while preserving throughput, since Cloud Functions scales instances independently of the memory setting.

Why this answer

The 'memory limit exceeded' error indicates that the function's allocated memory is insufficient for the workload during peak load. Increasing the memory allocation (Option B) directly resolves this by providing more RAM for processing larger messages or concurrent operations, without altering the invocation pattern or throughput. Cloud Functions (2nd gen) allow memory to be set up to 32 GiB, and this change does not reduce the number of events processed per second.

Exam trap

Google Cloud often tests the misconception that scaling out (more instances) solves memory issues, but the trap here is that memory limits are per-instance, so only increasing the per-instance memory allocation directly resolves the error.

How to eliminate wrong answers

Option A is wrong because increasing the maximum number of concurrent instances does not address the per-instance memory limit; it may actually worsen the problem by allowing more instances to hit the same memory ceiling simultaneously. Option C is wrong because Pub/Sub batching reduces the number of function invocations but does not increase the memory available per invocation; it could also increase latency and does not fix the root cause of memory exhaustion. Option D is wrong because splitting the function into multiple smaller functions does not increase the memory per function instance; it adds complexity and may reduce throughput due to additional overhead, without guaranteeing that each smaller function avoids memory limits.

26
MCQeasy

A company uses Cloud Spanner for a global financial application. They experience increased latency and transaction aborts during peak hours. Which measure should they take first to improve reliability?

A.Increase the number of nodes in the Spanner instance.
B.Reduce the number of indexes on frequently updated columns.
C.Optimize transactions to reduce lock contention.
D.Use interleaved tables to co-locate related data.
AnswerC

Reducing lock contention directly addresses the transaction aborts and latency spikes. Cloud Spanner aborts transactions when locks conflict, so shortening transactions and ordering reads and writes to minimise overlapping access lowers abort rates and improves throughput during peak load.

Why this answer

Transaction aborts and latency in Cloud Spanner are most commonly caused by lock contention during peak hours. By optimizing transactions—such as reducing their scope, using read-only transactions where possible, and avoiding hot-spot writes—you directly address the root cause of contention without incurring additional cost or schema changes. This aligns with Google's best practices for Spanner reliability.

Exam trap

Google Cloud often tests the misconception that scaling nodes (Option A) is the universal fix for performance issues, but the trap here is that Spanner's horizontal scaling does not resolve lock contention—it only increases parallelism, which can worsen contention if transactions are not optimized.

How to eliminate wrong answers

Option A is wrong because increasing nodes primarily improves throughput and storage capacity, not latency or abort rates caused by lock contention; adding nodes can even increase distributed transaction overhead. Option B is wrong because reducing indexes on frequently updated columns may reduce write amplification but does not address the immediate issue of lock contention and aborts; indexes are not the primary cause of transaction conflicts. Option D is wrong because interleaved tables co-locate parent-child rows for faster joins and lower latency, but they do not reduce lock contention; in fact, they can increase contention if the parent row becomes a hot spot.

27
MCQeasy

A developer wants to monitor a custom application metric from their application running on GKE. What should they use?

A.Cloud Logging
B.Cloud Trace
C.Cloud Debugger
D.Cloud Monitoring custom metrics API
AnswerD

Cloud Monitoring's custom metrics API accepts user-defined time series from GKE workloads, satisfying the requirement to monitor an application-specific metric rather than built-in system telemetry. The developer writes metric descriptors and time-series data via the API, which Cloud Monitoring then charts and alerts on.

Why this answer

Cloud Monitoring custom metrics API (option D) is the correct choice because it allows a developer to push custom application-specific metrics (e.g., request latency, queue depth) from a GKE pod using the `custom.googleapis.com` metric domain. This integrates directly with Cloud Monitoring for alerting and dashboards, whereas Cloud Logging is for log data, not metrics.

Exam trap

The trap here is that candidates confuse Cloud Logging (for logs) with Cloud Monitoring (for metrics), or assume that Cloud Trace can handle custom metrics because it deals with application performance data.

How to eliminate wrong answers

Option A is wrong because Cloud Logging ingests log entries (text-based events), not numeric metric data points; it cannot be used to monitor custom application metrics like counters or gauges. Option B is wrong because Cloud Trace is a distributed tracing system for latency analysis of requests, not for publishing custom numeric metrics. Option C is wrong because Cloud Debugger is used for inspecting application state at specific code points without stopping the app, not for collecting or monitoring time-series metrics.

28
MCQmedium

Your company runs a customer-facing API on Cloud Run with a concurrency setting of 80. The API calls a backend Cloud Function that performs a heavy computation (2–5 seconds). During peak hours, the API experiences increased latency and some requests time out after 60 seconds. Monitoring shows that the Cloud Run max instances is set to 100, and the Cloud Function max instances is set to 10. The timeout for Cloud Run is set to 300 seconds. The Cloud Function's timeout is set to 540 seconds. You need to reduce end-to-end latency and prevent timeouts while minimizing cost. Which action is most effective?

A.Increase Cloud Run max instances from 100 to 500
B.Increase Cloud Run request timeout from 300 to 600 seconds
C.Increase Cloud Function max instances from 10 to 100
D.Reduce Cloud Run concurrency from 80 to 10
AnswerC

Correct: removes backend capacity bottleneck.

Why this answer

The bottleneck is the Cloud Function's low max instances (10), causing queuing. Increasing Cloud Function max instances allows more concurrent requests to be processed, reducing latency and timeouts. Option A is wrong because concurrency on Cloud Run is separate from backend; reducing concurrency would require more Cloud Run containers and increase cost.

Option B is wrong because increasing Cloud Run max instances alone doesn't help if Cloud Function capacity is the limit. Option D is wrong because increasing Cloud Run timeout doesn't reduce latency; it just keeps the connection alive longer.

29
MCQmedium

A developer ran the above command to create a health check for a backend service. Which of the following should they do to resolve the error?

A.Change the request-path to a different value.
B.Delete the existing health check and recreate it.
C.Add the --global flag to the command.
D.Use --load-balancer-type internal to create a new health check with the same name.
E.Use a different name for the health check.
AnswerE

The error stems from a naming collision: a health check with that identifier already exists on the backend service. Supplying a unique name satisfies the API's uniqueness constraint, allowing the new health check to be created without altering its configuration.

Why this answer

The error indicates that a health check with the same name already exists. In Google Cloud, health check names must be unique within a project (or within a region for regional health checks). By using a different name, the developer can create a new health check without conflicting with the existing one.

Exam trap

Google Cloud often tests the misconception that modifying parameters like request-path or load balancer type can resolve naming conflicts, when in fact the core issue is a duplicate name that must be changed.

How to eliminate wrong answers

Option A is wrong because changing the request-path does not resolve a naming conflict; it only alters the path used for health checks. Option B is wrong because deleting and recreating the health check with the same name would still fail if the name is already in use. Option C is wrong because the --global flag is used for global accelerators, not for resolving health check naming conflicts.

Option D is wrong because --load-balancer-type internal specifies the load balancer type, not the health check name; it does not address the duplicate name error.

30
MCQhard

Your company runs a data pipeline on Google Cloud using Cloud Dataflow for streaming processing from Pub/Sub to BigQuery. The pipeline writes to a BigQuery table partitioned by day. The data is used for real-time dashboards. Recently, a spike in traffic caused the Dataflow pipeline to fall behind, and the dashboard displayed stale data. You need to design the pipeline to handle traffic spikes without data loss or long delays. The pipeline must be cost-efficient and use defaults where possible. Which solution should you implement?

A.Enable autoscaling in the Dataflow pipeline and use Streaming Engine to handle larger throughput
B.Modify the pipeline to use a batch (non-streaming) approach, writing hourly batches from Pub/Sub to BigQuery
C.Create a Cloud Scheduler job that increases the number of Dataflow workers every 5 minutes based on Pub/Sub subscription backlog
D.Change the Dataflow worker machine type from n1-standard-4 to n1-highmem-8
AnswerA

Autoscaling adds workers dynamically as Pub/Sub backlog grows, while Streaming Engine offloads pipeline state and shuffle processing to a managed service, raising throughput without the cost of permanently over-provisioned workers. Together they absorb traffic spikes with default settings, avoiding stale dashboards and data loss.

Why this answer

Enabling autoscaling in Dataflow allows the pipeline to dynamically adjust the number of workers based on the processing backlog, while Streaming Engine offloads the shuffle and state storage to Google-managed resources, reducing the impact of traffic spikes. This combination ensures the pipeline can scale up quickly to handle increased throughput without data loss or long delays, and it remains cost-efficient by scaling down when demand decreases.

Exam trap

Google Cloud often tests the misconception that manual scaling (Option C) or static resource changes (Option D) are sufficient for handling spikes, when in fact Dataflow's built-in autoscaling and Streaming Engine are the designed, cost-efficient solutions for dynamic workloads.

How to eliminate wrong answers

Option B is wrong because switching to a batch approach introduces inherent latency (hourly batches) that would make the real-time dashboard stale, violating the requirement for minimal delays; it also does not handle spikes within the batch window. Option C is wrong because using Cloud Scheduler to manually adjust worker count every 5 minutes is reactive, not adaptive, and cannot respond quickly enough to sudden spikes; Dataflow's native autoscaling is designed to adjust more granularly and efficiently. Option D is wrong because simply changing the worker machine type to a larger instance (n1-highmem-8) does not address the need for dynamic scaling; it increases cost without guaranteeing sufficient capacity during spikes and does not leverage Dataflow's autoscaling capabilities.

31
MCQmedium

The exhibit shows the output of a 'gcloud compute instances describe' command for an instance. What is the most likely impact on reliability if the host machine needs maintenance?

A.The instance will be terminated and then restarted, causing a brief downtime.
B.The instance will not be affected because automatic restart is enabled.
C.The instance will be backed up automatically before maintenance.
D.The instance will be live migrated to another host without interruption.
AnswerA

With TERMINATE, the instance is shut down and later restarted on a healthy host, resulting in downtime.

Why this answer

When a host machine requires maintenance, Google Compute Engine instances that are not configured for live migration will be terminated and then restarted on another host. This behavior is determined by the 'onHostMaintenance' setting; if it is set to 'TERMINATE' (the default for instances with GPUs or preemptible VMs), the instance stops and restarts, causing brief downtime. The exhibit likely shows 'onHostMaintenance: TERMINATE' or the instance lacks live migration support, making termination and restart the expected outcome.

Exam trap

Google Cloud often tests the distinction between 'automatic restart' (which handles crash recovery) and 'onHostMaintenance' (which handles planned maintenance), causing candidates to mistakenly think automatic restart prevents downtime during maintenance.

How to eliminate wrong answers

Option B is wrong because 'automatic restart' is a separate setting that controls whether an instance restarts after a failure or crash, not how it behaves during host maintenance; it does not prevent downtime from maintenance events. Option C is wrong because Google Compute Engine does not automatically back up instances before host maintenance; backups must be configured separately via snapshots or images. Option D is wrong because live migration is only possible if the instance has 'onHostMaintenance' set to 'MIGRATE' and does not have GPUs, local SSDs, or preemptible status; the exhibit likely shows a configuration that disables live migration, such as a GPU attached or the setting explicitly set to 'TERMINATE'.

32
MCQmedium

A company runs a web application on Google Kubernetes Engine (GKE) with Cluster Autoscaler enabled. During a traffic spike, the application becomes slow and some requests timeout. The cluster has sufficient CPU and memory headroom. What is the most likely cause and solution?

A.Increase the node pool's machine type to a larger size.
B.Enable Cluster Autoscaler to add more nodes.
C.Deploy the application in a regional cluster for higher availability.
D.Configure Horizontal Pod Autoscaler (HPA) based on CPU utilization or custom metrics.
AnswerD

Cluster Autoscaler only adds nodes when pods are pending; with CPU and memory headroom already available, no new nodes are needed. The bottleneck is pod count, so Horizontal Pod Autoscaler must scale replicas based on CPU or custom metrics to absorb the traffic spike.

Why this answer

The cluster has sufficient CPU and memory headroom, indicating that the issue is not about cluster capacity but about pod-level scaling. The Horizontal Pod Autoscaler (HPA) automatically scales the number of pod replicas based on observed CPU utilization or custom metrics, which directly addresses the application slowdown and timeouts during traffic spikes by distributing the load across more pods.

Exam trap

Google Cloud often tests the distinction between node-level scaling (Cluster Autoscaler) and pod-level scaling (HPA), trapping candidates who assume that adding more nodes is the solution when the cluster already has headroom, whereas the real issue is insufficient pod replicas to handle the load.

How to eliminate wrong answers

Option A is wrong because increasing the node pool's machine type addresses node-level resource constraints, but the cluster already has sufficient CPU and memory headroom, so the bottleneck is at the pod level, not the node level. Option B is wrong because Cluster Autoscaler is already enabled and the cluster has headroom, so adding more nodes would not solve the problem of insufficient pod replicas to handle the traffic spike. Option C is wrong because deploying in a regional cluster improves availability and resilience to zone failures, but does not directly address the performance degradation and timeouts caused by insufficient application instances during a traffic spike.

33
Multi-Selecthard

A healthcare analytics company runs a stateless API on a regional managed instance group behind an external Application Load Balancer. The SRE team wants to improve reliability and reduce customer-visible errors during zonal and instance failures. (Choose two.)

Select 2 answers
A.Set the load balancer's balancing mode to RATE and configure a maximum rate per instance to cap traffic.
B.Configure the load balancer's backend service with a health check that matches the API's readiness endpoint and set a sensible unhealthy threshold.
C.Enable autoscaling on the managed instance group based on CPU utilization and load balancing capacity.
D.Create a second managed instance group in a different region and add it as a backend to the same load balancer.
E.Enable Cloud CDN on the backend service to cache API responses and reduce origin load.
AnswersB, C

An accurate health check lets the load balancer stop sending traffic to instances that cannot serve requests, so users are routed only to healthy backends. Matching the readiness endpoint ensures the check reflects real application health rather than just port availability. This reduces customer-visible errors during instance failures and is a core reliability practice for load-balanced stateless services.

Why this answer

Autoscaling lets the regional managed instance group add or maintain capacity so healthy instances in surviving zones can handle traffic when one zone or instance fails. A health check aligned to the API's readiness endpoint ensures the load balancer routes only to instances that can actually serve requests, cutting customer-visible errors. Together they address zonal and instance failure reliability.

Exam trap

The trap here is reaching for regional expansion or caching when the stated failure domain is zonal and instance-level, where autoscaling and accurate health checking are the targeted fixes.

34
MCQmedium

Refer to the exhibit. An application running on a GCE instance (ID: 1234567890) is unable to connect to a database at 10.0.0.1:5432. The logs show repeated 'Connection refused' errors. What is the most likely cause?

A.The firewall rule allowing traffic on port 5432 is missing or misconfigured.
B.The instance is using an outdated SSL certificate.
C.The database service is not running or is not listening on port 5432.
D.The VPC network has no route to the database subnet.
AnswerC

A refused TCP connection means the host is reachable but nothing accepts traffic on port 5432, indicating the database process is stopped or bound elsewhere. Firewall or routing problems would typically produce timeouts, not refusals.

Why this answer

The 'Connection refused' error indicates that the TCP handshake was rejected by the target host, which typically means the database service is not actively listening on port 5432. This is distinct from a firewall block, which would result in a timeout or 'no route to host' error. Since the error is immediate and specific to port 5432, the most likely cause is that the PostgreSQL or other database service is not running or is bound to a different interface/port.

Exam trap

Google PCA exams often test the distinction between firewall blocks (timeout) and service unavailability (connection refused), so the trap here is that candidates confuse a missing firewall rule with a service not listening, even though the error messages are fundamentally different.

How to eliminate wrong answers

Option A is wrong because a missing or misconfigured firewall rule would cause a timeout or 'connection timed out' error, not an immediate 'Connection refused' — the latter requires the host to actively reject the connection. Option B is wrong because SSL certificate issues would manifest as TLS handshake failures or certificate validation errors, not a raw TCP-level 'Connection refused'. Option D is wrong because if there were no route to the database subnet, the error would be 'No route to host' or a network unreachable message, not a port-specific refusal.

35
MCQeasy

A retail company runs a public-facing API on Cloud Run in the europe-west1 region. During a marketing campaign, traffic tripled within minutes. The service remained available, but some requests returned HTTP 503 errors. The team wants to reduce the chance of 503 errors during future traffic spikes while keeping the deployment simple. What should they do?

A.Set the Cloud Run service's maximum number of instances to a high value and configure a minimum number of instances greater than zero.
B.Increase the container's allocated memory and CPU limits in the Cloud Run revision settings.
C.Enable Cloud CDN on the Cloud Run service and set a long cache TTL for all responses.
D.Deploy the service to multiple Cloud Run regions and use a global external Application Load Balancer with a serverless network endpoint group.
AnswerA

Cloud Run scales out by adding instances, but if the configured maximum instance count is reached, additional requests can be rejected with 503 errors. Raising the maximum allows more instances to handle the spike, and setting a minimum above zero keeps warm instances ready so cold starts do not contribute to failures during the initial surge. This directly addresses the observed 503 behavior under burst load.

Why this answer

The 503 errors occurred because the service hit its configured maximum instance count during the spike. Raising the maximum instance limit lets Cloud Run create enough instances to absorb the burst, and setting a minimum instance count above zero keeps warm instances available so the initial surge does not fail while new instances start. Resource limits, CDN caching, and multi-region deployment do not directly remove the instance ceiling that caused the rejections.

Exam trap

The trap here is assuming that increasing CPU or memory on a Cloud Run service increases its capacity to handle concurrent requests, when the real constraint during a burst is the maximum instance count.

36
MCQmedium

The exhibit shows a Cloud Storage bucket configuration. What does this configuration ensure?

A.Older versions of objects are automatically transferred to a different storage class.
B.Data is replicated to another region for disaster recovery.
C.Objects can only be permanently deleted after the retention period expires.
D.Objects older than 30 days will be automatically deleted.
AnswerC

A bucket retention policy with a set retention period locks each object until that period elapses, so deletion requests are refused beforehand. This directly enforces the configuration's guarantee that permanent deletion is possible only once the retention period expires.

Why this answer

The exhibit shows a bucket configured with a retention policy. When a retention policy is set on a Cloud Storage bucket, objects cannot be deleted or overwritten until the retention period expires. This ensures that objects can only be permanently deleted after the retention period ends, which is exactly what option C describes.

Exam trap

The trap here is that candidates confuse retention policies with lifecycle management rules, mistakenly thinking retention policies automatically delete or transition objects, when in fact they only prevent deletion until the retention period expires.

How to eliminate wrong answers

Option A is wrong because retention policies do not automatically transfer objects to a different storage class; that is the function of lifecycle management rules, not retention policies. Option B is wrong because retention policies do not replicate data to another region; replication is configured separately using object replication or dual-region buckets. Option D is wrong because retention policies do not automatically delete objects after a period; they prevent deletion until the retention period expires, and automatic deletion is achieved via lifecycle rules with a Delete action.

37
MCQhard

Refer to the exhibit. A Cloud Deploy pipeline has a release with two targets: staging and prod. The staging rollout succeeded, but the prod rollout failed with 'MANIFEST_INVALID'. What is the most likely cause of the failure?

A.The manifest for prod contains a syntax error or references a resource that does not exist in the prod cluster.
B.The prod target's applyManifest has a higher replica count than staging, which violates a cluster quota.
C.The prod cluster does not have the necessary permissions to pull the container image.
D.The release was not approved for the prod target.
AnswerA

A manifest that parses cleanly against staging can still fail validation against prod, because MANIFEST_INVALID is raised when the rendered manifest is syntactically malformed or references resources absent from the target cluster. The prod target's differing namespace, CRDs or API versions therefore satisfy the stem's constraint that staging succeeded while prod alone failed.

Why this answer

The 'MANIFEST_INVALID' error in Cloud Deploy indicates that the Kubernetes manifest provided for the prod target is syntactically incorrect or references a resource (e.g., a ConfigMap, Secret, or custom resource definition) that does not exist in the prod cluster. This is a validation failure that occurs before any deployment attempt, so it is not related to runtime issues like permissions or quotas.

Exam trap

A common trap is confusing 'MANIFEST_INVALID' with permission or quota issues, which are separate failure modes in the deployment pipeline.

How to eliminate wrong answers

Option B is wrong because a replica count exceeding a cluster quota would produce a different error, such as 'QUOTA_EXCEEDED' or a resource allocation failure during rollout, not a manifest validation error. Option C is wrong because insufficient permissions to pull a container image would result in an 'ImagePullBackOff' or 'ErrImagePull' error at the pod level, not a manifest validation error during the deploy step. Option D is wrong because a missing approval for the prod target would cause the rollout to be pending or skipped, not to fail with 'MANIFEST_INVALID'; approval gates are checked before the rollout begins.

38
MCQmedium

A company deploys a microservices application on Google Kubernetes Engine (GKE). Pods in one deployment are frequently OOMKilled. The team sets memory requests and limits, but pods still crash. What is the most likely remaining cause?

A.CPU requests are too low, causing throttling and eventual crash.
B.The node pool is too small, causing memory pressure on the node.
C.Memory limits are set higher than the node's allocatable memory.
D.The application has a memory leak that eventually exceeds the limit.
AnswerD

Requests and limits only cap consumption; they cannot prevent a leak from growing until the container exceeds its limit and is OOMKilled. Since configuration is already correct, the remaining cause is application-level: unbounded allocation that eventually surpasses the configured memory limit.

Why this answer

OOMKilled errors occur when a container exceeds its memory limit. Setting memory requests and limits prevents unbounded usage, but if the application has a memory leak, it will continue to consume memory until it hits the configured limit, causing the kernel's Out-Of-Memory (OOM) killer to terminate the pod. The fact that pods still crash after setting limits indicates the application itself is the root cause, not resource configuration.

Exam trap

The trap here is that candidates confuse OOMKilled (per-container limit) with node-pressure eviction (node-level memory), or assume that setting requests/limits automatically fixes all memory issues, ignoring application-level bugs like memory leaks.

How to eliminate wrong answers

Option A is wrong because CPU throttling does not cause OOMKilled; CPU limits throttle performance but do not trigger the OOM killer, which is specific to memory exhaustion. Option B is wrong because node-level memory pressure would cause pods to be evicted (not OOMKilled) or the node to become NotReady, but the question states pods are OOMKilled, which is a per-container limit violation, not a node-level issue. Option C is wrong because setting memory limits higher than the node's allocatable memory would prevent the pod from being scheduled (pending state), not cause it to run and then be OOMKilled.

39
MCQhard

A financial services company is migrating a monolithic Java application to Google Kubernetes Engine (GKE) for improved scalability and reliability. The application serves real-time trading data and has strict latency requirements. Post-migration, the team observes frequent pod restarts due to OutOfMemory (OOM) errors, increased latency during peak trading hours, and occasional database connection timeouts. The current setup uses a single GKE cluster with a node pool of n1-standard-4 machines, a stateless application deployed as a Deployment with resource requests and limits set to 512 Mi memory and 1 CPU. The database is a Cloud SQL PostgreSQL instance with 2 vCPUs and 7.5 GB memory, and applications connect using a hardcoded connection string. The team wants to ensure reliable operation under load and during node maintenance events. Which course of action best addresses the reliability issues?

A.Adjust resource requests to 1 Gi memory and 2 CPU, set limits to 2 Gi and 4 CPU, create an HPA based on a custom metric (e.g., requests per second), enable cluster autoscaler, implement Cloud SQL connection pooling via Cloud SQL Auth Proxy with a max connection pool size, and configure PDB with maxUnavailable 1.
B.Enable GKE node auto-upgrade, configure Pod Disruption Budgets (PDB) with minAvailable 1, and set readiness probes to check application health.
C.Migrate the database to a StatefulSet in GKE with persistent volumes, increase node count to 10, and enable cluster autoscaler.
D.Increase memory limits to 2 Gi and CPU to 2, add Horizontal Pod Autoscaler (HPA) based on CPU utilization, and implement connection pooling using Cloud SQL Auth Proxy.
AnswerA

Raising memory requests and limits above the current 512 Mi stops OOM kills, while the HPA, cluster autoscaler, Cloud SQL Auth Proxy pooling and PDB together address peak-load latency, connection exhaustion and node maintenance disruption — the four reliability symptoms named in the stem.

Why this answer

Best addresses all reliability issues. Adjusting resource requests to 1 Gi memory and 2 CPU ensures proper scheduling, while limits of 2 Gi and 4 CPU prevent OOM errors. The HPA based on custom metrics (e.g., requests per second) scales pods proactively during peak trading hours.

Cluster autoscaler handles node capacity, and Cloud SQL connection pooling via Cloud SQL Auth Proxy with a max pool size prevents database connection timeouts. Finally, a PDB with maxUnavailable 1 ensures availability during node maintenance. Option B misses resource tuning, autoscaling, and connection pooling.

Option C unnecessarily moves the database to GKE, increasing complexity and losing managed DB benefits. Option D lacks custom metric HPA, cluster autoscaler, and PDB, leaving gaps in scalability and maintenance handling.

40
MCQmedium

Your team manages a service with a 99.9% uptime SLO over a 30-day window. The error budget for this period is 43 minutes. In the first week, outages consumed 30 minutes of the budget. You are planning a new release. What should you do?

A.Reduce the SLO to 99.8% to increase the error budget.
B.Proceed with the release because the remaining budget is sufficient.
C.Delay the release and focus on improving reliability to rebuild the error budget.
D.Release the feature but only to a small percentage of users.
AnswerC

Conservative approach: wait until more error budget is earned (e.g., through flawless operation) before releasing.

Why this answer

With only 13 minutes of error budget remaining after the first week, proceeding with the release (Option B) risks exhausting the budget entirely from any unforeseen issues, violating the 99.9% SLO. Delaying the release (Option C) allows the team to focus on reliability improvements, such as implementing canary deployments, adding circuit breakers, or enhancing monitoring with tools like Prometheus and Grafana, to rebuild the error budget over the remaining 23 days. This aligns with the principle of using error budgets to balance innovation with reliability, as defined in Google's SRE practices.

Exam trap

Google Cloud often tests the misconception that a canary release (Option D) is always safe, but the trap here is that it still consumes error budget and does not solve the underlying reliability deficit when the budget is already critically low.

How to eliminate wrong answers

Option A is wrong because reducing the SLO to 99.8% would increase the error budget to 86.4 minutes, but this is a reactive measure that lowers the reliability target rather than addressing the root cause of the outages; it also violates the principle of maintaining a consistent SLO commitment to customers. Option B is wrong because proceeding with the release with only 13 minutes of error budget left is reckless—any minor incident could exhaust the budget, leading to SLO violations and potential service credits or customer dissatisfaction, especially since the first week already consumed 70% of the budget. Option D is wrong because releasing to a small percentage of users (e.g., a canary deployment) is a valid risk mitigation strategy, but it does not address the fact that the error budget is nearly depleted; even a small-scale release could introduce bugs that consume the remaining budget, and the team should first stabilize the service before any new changes.

41
MCQeasy

A company deploys a web application on Compute Engine behind an HTTP Load Balancer. They want to ensure only healthy instances receive traffic. What should they configure?

A.Configure the instance group autoscaling based on CPU utilization
B.Configure an HTTP health check with a custom request path that returns a 200 status
C.Configure a TCP health check on port 80
D.Configure an SSL health check to verify TLS handshake
AnswerB

An HTTP health check probes a specified path and marks an instance healthy only when it returns HTTP 200, so the load balancer routes traffic exclusively to instances passing that check. This directly satisfies the requirement that only healthy instances receive traffic.

Why this answer

An HTTP health check with a custom request path that returns a 200 status allows the HTTP Load Balancer to verify that the web application is actually serving requests correctly. This ensures that only instances passing the application-level health check are considered healthy and receive traffic, preventing requests from being routed to instances that may be running but not serving the expected content.

Exam trap

The trap here is that candidates often confuse health checks with autoscaling metrics, assuming that CPU-based autoscaling alone ensures traffic is only sent to healthy instances, when in fact health checks are a separate mechanism required for load balancer traffic routing.

How to eliminate wrong answers

Option A is wrong because autoscaling based on CPU utilization manages the number of instances but does not determine which instances are healthy for traffic routing; the load balancer still needs health checks to decide which instances to send traffic to. Option C is wrong because a TCP health check on port 80 only verifies that the TCP port is open, not that the web application is responding correctly; an instance could have a listening port but return errors or be unresponsive at the application layer. Option D is wrong because an SSL health check verifies the TLS handshake, which is unnecessary for HTTP traffic and does not validate the application's response; it is designed for HTTPS backends, not plain HTTP.

42
Multi-Selectmedium

A company wants to improve the reliability of their microservices architecture on Google Cloud. Which TWO practices should they implement? (Choose 2)

Select 2 answers
A.Design with a single point of failure for simplicity
B.Implement retry with exponential backoff
C.Use synchronous communication between all services
D.Implement circuit breaker pattern
E.Disable health checks to reduce latency
AnswersB, D

Retry with backoff handles transient failures without overwhelming the system.

Why this answer

B is correct because implementing retry with exponential backoff allows transient failures (e.g., network timeouts, temporary service unavailability) to be handled gracefully by automatically retrying the request after increasing delays, reducing load on the recovering service. This pattern is essential in microservices on Google Cloud to improve reliability without overwhelming downstream dependencies.

Exam trap

Google Cloud often tests the misconception that synchronous communication is more reliable because it provides immediate feedback, but in distributed systems, asynchronous patterns and resilience mechanisms like retries and circuit breakers are actually critical for reliability.

43
Multi-Selecthard

A company runs a microservices-based application on Google Kubernetes Engine (GKE) with a Regional cluster. They want to improve reliability by implementing best practices for pod scheduling and resilience. Which TWO actions should they take? (Choose two.)

Select 2 answers
A.Set terminationGracePeriodSeconds to 0 for faster pod termination during scale-down
B.Enable cluster autoscaler to automatically add nodes when pods are pending
C.Define a PodDisruptionBudget for each deployment to limit the number of concurrent disruptions
D.Set resource requests equal to limits to ensure guaranteed QoS class
E.Configure pod anti-affinity to spread replicas across different zones
AnswersC, E

Correct: PDB ensures minimum availability during voluntary disruptions.

Why this answer

A PodDisruptionBudget (PDB) limits the number of Pods of a replicated application that can be down simultaneously from voluntary disruptions, such as node maintenance or cluster upgrades. This ensures that a minimum number of replicas remain available, improving application reliability during planned events.

Exam trap

Google Cloud often tests the distinction between voluntary disruptions (handled by PDB) and involuntary disruptions (e.g., node failure), and the trap here is that candidates confuse resource optimization (requests/limits) or scaling (cluster autoscaler) with resilience mechanisms like PDB and anti-affinity.

44
MCQhard

A company runs a stateful application on Compute Engine with persistent disks. They want to ensure data durability across a zone failure. What is the best approach?

A.Replicate data at application level to another instance in a different zone
B.Use Google Cloud NetApp Volumes with replication
C.Use regional persistent disks
D.Take regular snapshots of the persistent disks and store them in a multiregional bucket
AnswerC

Regional persistent disks synchronously replicate data across two zones within the same region, so a zone failure leaves a healthy replica available. Zonal persistent disks reside in a single zone and cannot satisfy the cross-zone durability constraint.

Why this answer

Regional persistent disks (RPDs) synchronously replicate data between two zones in the same region, providing an RPO of zero and automatic failover without application-level changes. This ensures data durability across a zone failure while maintaining consistent performance and low latency.

Exam trap

Google Cloud often tests the distinction between synchronous replication (regional persistent disks) and asynchronous backup (snapshots), leading candidates to choose snapshots for durability when they actually need zero RPO across a zone failure.

How to eliminate wrong answers

Option A is wrong because replicating data at the application level adds complexity, latency, and requires custom code, whereas Compute Engine offers a managed, synchronous replication solution. Option B is wrong because Google Cloud NetApp Volumes is a third-party service that is not natively integrated with Compute Engine for this use case and introduces additional cost and management overhead. Option D is wrong because regular snapshots stored in a multiregional bucket provide point-in-time recovery but have an RPO of minutes to hours and do not offer synchronous replication, so data written between snapshots is lost during a zone failure.

45
Multi-Selectmedium

Your organization is implementing a Disaster Recovery plan for a critical database. Which THREE components are essential for a robust DR strategy? (Choose 3)

Select 3 answers
A.A single global load balancer for both regions.
B.Automated failover process to switch traffic to the DR region.
C.Data replication strategy (synchronous or asynchronous) to a secondary region.
D.Regular DR drills (testing failover at least once per quarter).
E.Using a single zone for the primary region.
AnswersB, C, D

Automation minimizes manual errors and reduces RTO.

Why this answer

An automated failover process is essential for minimizing Recovery Time Objective (RTO) in a Disaster Recovery strategy. Without automation, manual intervention introduces delays and risks of human error, which can extend downtime significantly. In cloud or on-premises environments, automated failover typically relies on health checks, DNS updates, or traffic manager rules to seamlessly redirect traffic to the DR region when the primary fails.

Exam trap

Google Cloud often tests the misconception that a single global load balancer provides high availability, when in fact it becomes a single point of failure unless it is itself deployed in a redundant, multi-region architecture.

46
MCQhard

An organization is migrating a legacy monolithic application to Google Cloud. The application currently runs on a single server with an on-premises database. The application is stateful and requires low-latency access to the database. The migration must minimize downtime and ensure high availability. Which architecture should the company adopt?

A.Deploy on GKE with StatefulSets and use Cloud Spanner for global consistency.
B.Deploy on Compute Engine with a regional persistent disk and use Cloud SQL for PostgreSQL with regional high availability.
C.Deploy on App Engine Standard Environment and use Cloud Firestore in Datastore mode.
D.Deploy on Cloud Run and use Cloud SQL with read replicas.
AnswerB

A regional persistent disk keeps the stateful application's data synchronously replicated across zones, while Cloud SQL for PostgreSQL regional HA provides an automatic standby in a second zone. Together they deliver the required high availability and low-latency database access, and minimise downtime during cutover.

Why this answer

It combines Compute Engine with a regional persistent disk for synchronous replication across zones, ensuring high availability with minimal downtime during a zonal failure. Cloud SQL for PostgreSQL with regional high availability provides a managed, low-latency database with automatic failover, meeting the stateful application's need for low-latency access and high availability without the complexity of container orchestration.

Exam trap

The trap here is that candidates often overcomplicate the solution by choosing containerized or serverless options (GKE, Cloud Run, App Engine) without recognizing that a legacy monolithic stateful application with low-latency requirements is best served by a simple, proven VM-based architecture with regional persistent disks and a managed relational database with synchronous replication.

How to eliminate wrong answers

Option A is wrong because GKE with StatefulSets introduces orchestration overhead and potential downtime during cluster upgrades or node failures, and Cloud Spanner, while globally consistent, adds latency and cost overkill for a single-region low-latency requirement. Option C is wrong because App Engine Standard Environment is stateless by design and does not support stateful applications with persistent local storage, and Cloud Firestore in Datastore mode is a NoSQL database that does not provide the relational consistency and low-latency access expected from a legacy monolithic database. Option D is wrong because Cloud Run is stateless and ephemeral, requiring external storage for state, and Cloud SQL with read replicas does not provide synchronous replication for high availability; read replicas are asynchronous and cannot guarantee zero data loss during a failover.

47
MCQmedium

A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. During a traffic spike, some instances become unhealthy but are not automatically replaced. What is the most likely cause?

A.The MIG is regional and one zone failed.
B.The autohealing health check is misconfigured.
C.The instance template has a startup script error.
D.The HTTP load balancer's health check is failing.
AnswerB

Autohealing relies on a health check to detect and recreate unhealthy instances. If that health check is misconfigured — wrong path, port, or thresholds — the MIG never marks instances unhealthy, so no automatic replacement occurs despite the traffic spike.

Why this answer

The most likely cause is that the autohealing health check is misconfigured. In a managed instance group, autohealing relies on a health check to detect unhealthy instances and trigger replacement. If the health check is misconfigured (e.g., wrong port, path, or protocol), the MIG will not recognize instances as unhealthy and will not automatically replace them, even during a traffic spike.

Exam trap

Google Cloud often tests the distinction between the MIG's autohealing health check and the load balancer's health check, leading candidates to incorrectly attribute instance replacement failures to load balancer issues rather than the MIG's own health check configuration.

How to eliminate wrong answers

Option A is wrong because a regional MIG with a single zone failure would still trigger autohealing in the remaining healthy zones, and the MIG would replace instances in the failed zone if the health check is correctly configured. Option C is wrong because a startup script error would cause instances to fail at boot, but the MIG would still attempt to replace them based on the health check; the issue is not about the template but the detection mechanism. Option D is wrong because the HTTP load balancer's health check is separate from the MIG's autohealing health check; a failing load balancer health check does not prevent the MIG from replacing unhealthy instances if its own health check is properly configured.

48
Multi-Selecthard

Which THREE options are valid strategies for disaster recovery (DR) in Google Cloud?

Select 3 answers
A.Store hourly snapshots of Compute Engine disks in the same region.
B.Deploy a mirrored environment in another region and use Traffic Director to fail over.
C.Enable Cloud CDN to cache static content from multiple origins.
D.Use a Cloud Storage bucket in a different region with Object Versioning enabled.
E.Configure a cross-region replica for Cloud SQL and promote it during failover.
AnswersB, D, E

A mirrored environment in a second region provides a warm or hot standby, and Traffic Director performs global load balancing with health-checked failover, redirecting traffic when the primary region fails. This satisfies the requirement for a cross-region DR strategy on Google Cloud.

Why this answer

Option B is correct because deploying a mirrored environment in a second region and using Traffic Director for global load balancing and failover provides true cross-region DR, ensuring workloads survive a full regional outage. Option D is correct because a Cloud Storage bucket in a different region with Object Versioning enabled gives durable, geographically separated copies of data and protects against accidental deletion or overwrite, which is a valid DR strategy. Option E is correct because a Cloud SQL cross-region replica can be promoted to a standalone primary during a regional failure, restoring database availability in another region.

Option A is not a valid DR strategy because snapshots stored in the same region are lost if that region fails, so they do not provide disaster recovery. Option C is not a DR strategy because Cloud CDN caching static content from multiple origins improves performance and availability but does not by itself recover from a regional disaster or protect stateful data.

Exam trap

The trap here is confusing high-availability features (like snapshots or CDN) with true disaster recovery, which requires geographic separation and automated failover mechanisms.

49
MCQeasy

A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. Users report intermittent 503 errors during traffic spikes. Which action should the company take to improve reliability?

A.Change the load balancer from regional to global
B.Configure a health check with a sufficient initial delay (grace period) in the MIG
C.Increase the autoscaling cool-down period from 60s to 120s
D.Increase the maximum number of instances in the MIG
AnswerB

During autoscaling spikes, new instances need time to boot before serving. A health check with a sufficient initial delay prevents the MIG from marking them unhealthy and removing them prematurely, eliminating the 503 errors caused by premature traffic routing.

Why this answer

Intermittent 503 errors during traffic spikes often indicate that new VM instances are being started but are not yet ready to serve traffic, causing the load balancer to forward requests to them prematurely. Configuring a health check with a sufficient initial delay (grace period) in the MIG ensures that newly created instances are given time to fully initialize and pass health checks before they receive traffic, preventing 503 errors. This directly addresses the root cause by allowing the application to become healthy before being added to the load balancer's backend.

Exam trap

Google Cloud often tests the misconception that scaling-related errors are always solved by increasing capacity or adjusting scaling parameters, when in fact the root cause is often a misconfigured health check or insufficient initialization time for new instances.

How to eliminate wrong answers

Option A is wrong because changing the load balancer from regional to global does not address the timing issue of new instances being marked healthy before they are ready; global load balancers improve cross-region routing but do not affect instance readiness. Option C is wrong because increasing the autoscaling cool-down period from 60s to 120s only delays the scaling decision after a scale-out event, but does not prevent the load balancer from sending traffic to instances that are still initializing; the cool-down period controls how often autoscaler evaluates metrics, not instance readiness. Option D is wrong because increasing the maximum number of instances in the MIG allows more capacity but does not fix the problem of instances being added to the backend pool before they are ready; it may even exacerbate the issue by creating more unhealthy instances.

50
MCQhard

You are responsible for incident management for a production service. You want to reduce manual toil during the initial response to common issues like high latency. What is the best approach?

A.Use Cloud Monitoring to trigger a Cloud Function that performs automated checks and rolls back the last deployment if latency spikes.
B.Set up Cloud Monitoring alerts with email notifications to the on-call engineer.
C.Create detailed runbooks and require the on-call to follow them step by step.
D.Enable Cloud Logging and set up a custom dashboard for the on-call.
AnswerA

Cloud Monitoring alerting policies detect the latency spike and invoke a Cloud Function that runs automated diagnostics and rolls back the last deployment, removing manual toil from initial response. This closed-loop remediation satisfies the stem's requirement to automate first response to common incidents.

Why this answer

It directly reduces manual toil by automating the initial response to common issues like high latency. Cloud Monitoring triggers a Cloud Function that performs automated checks and, if latency spikes, rolls back the last deployment, eliminating the need for human intervention during the critical first response phase.

Exam trap

Google Cloud often tests the distinction between 'alerting' (which still requires manual action) and 'automated remediation' (which reduces toil), so candidates mistakenly choose options that provide visibility or documentation instead of automation.

How to eliminate wrong answers

Option B is wrong because email notifications alone still require the on-call engineer to manually investigate and respond, which does not reduce toil; it merely alerts them. Option C is wrong because requiring the on-call to follow runbooks step by step still involves manual effort and does not automate the response, leaving toil unchanged. Option D is wrong because enabling Cloud Logging and setting up a custom dashboard provides visibility but does not automate any action, so the on-call must still manually diagnose and respond to the issue.

51
MCQmedium

Your organization uses Cloud Spanner for a customer database with a 99.999% availability SLA. You need a Disaster Recovery plan that ensures data consistency with zero RPO in case of a region failure. What should you do?

A.Use a single-region instance configuration and enable read replicas.
B.Export the database periodically to Cloud Storage and set up a cross-region load balancer.
C.Configure daily backups and store them in Cloud Storage in a different region.
D.Use a multi-region instance configuration (e.g., nam-eur-asia) for the Spanner instance.
AnswerD

A multi-region instance configuration synchronously replicates data across regions using Paxos consensus, so committed writes survive a regional failure with zero data loss. This directly satisfies the stem's zero RPO and consistency requirements, and the 99.999% SLA is only achievable with multi-region configurations, not regional ones.

Why this answer

Cloud Spanner multi-region instance configurations (e.g., nam-eur-asia) provide synchronous replication across multiple regions, ensuring strong global consistency and zero RPO. This architecture uses Paxos-based replication to commit writes only after they are durably stored in a majority of regions, so a region failure does not lose any committed data. The 99.999% availability SLA is met by automatic failover within the multi-region setup without manual intervention.

Exam trap

Google Cloud often tests the misconception that read replicas or periodic exports can achieve zero RPO, but only synchronous multi-region replication (as in Spanner's multi-region configurations) guarantees no data loss during a region failure.

How to eliminate wrong answers

Option A is wrong because single-region instance configurations with read replicas are not supported in Cloud Spanner; Spanner uses writable replicas, not read replicas, and a single-region setup cannot survive a full region failure, thus cannot achieve zero RPO. Option B is wrong because exporting the database periodically to Cloud Storage introduces a non-zero RPO (the time between exports) and does not guarantee data consistency at the point of failure; cross-region load balancers do not handle Spanner's transactional consistency. Option C is wrong because daily backups stored in a different region provide point-in-time recovery with a minimum RPO of 24 hours (or more), not zero RPO, and cannot ensure data consistency for transactions in flight at the time of failure.

52
MCQmedium

A retail company runs a customer-facing API on a managed instance group (MIG) of Compute Engine VMs behind an external Application Load Balancer. The SRE team wants the load balancer to stop sending traffic to a VM as soon as the local application health endpoint starts returning HTTP 500, even before the VM is fully unresponsive. Which load balancer component must be configured to achieve this?

A.Set the MIG's autoscaling policy to scale in when CPU utilization drops below a target threshold.
B.Enable Cloud CDN on the backend service so that failing responses are cached and served from edge locations.
C.Create an uptime check in Cloud Monitoring against the API endpoint and rely on its alerting policy to failover.
D.Configure an HTTP health check with a check interval and unhealthy threshold on the backend service, and attach it to the MIG.
AnswerD

The backend service's health check is exactly what drives traffic steering for the external Application Load Balancer. Pointing the health check at the application's health endpoint means a VM returning HTTP 500 is marked unhealthy after the unhealthy threshold is reached, so the load balancer drains it from rotation while the VM is still running.

Why this answer

For an external Application Load Balancer, the backend service's health check is the mechanism that determines which instances receive traffic. By pointing an HTTP health check at the application's health endpoint, a VM that begins returning HTTP 500 is marked unhealthy once the unhealthy threshold is met and is removed from rotation, so users are not routed to a failing instance while it is still alive.

Exam trap

The trap here is assuming monitoring uptime checks or autoscaling policies steer load balancer traffic, when only the backend service health check actually controls instance rotation.

53
MCQmedium

You are responsible for a Cloud Run service that experiences occasional cold starts, causing increased latency. You want to minimize cold starts while keeping costs under control. What should you do?

A.Set the minimum number of instances to a value greater than zero.
B.Increase the maximum number of instances.
C.Use a larger container image.
D.Deploy the service with a higher CPU limit.
AnswerA

Setting the minimum number of instances ensures that at least that many instances are always running, eliminating cold starts for the baseline traffic. You can set it to a low value, such as 1 or 2, to balance cost and performance. This is the recommended approach to reduce cold starts while controlling costs.

Why this answer

To minimize cold starts in Cloud Run, set the minimum number of instances to a value greater than zero. This keeps a baseline of warm instances ready to serve requests, reducing latency. Increasing the maximum instances or CPU limit does not prevent cold starts, and a larger container image would exacerbate them.

Exam trap

The trap here is thinking that increasing the maximum instances or CPU will prevent cold starts, but cold starts are only mitigated by keeping instances warm via minimum instances.

54
MCQhard

You are investigating a Vertex AI Workbench instance (instance-2) that is showing UNHEALTHY status. Based on the exhibit, what is the most likely cause of the issue?

A.The container image gcr.io/my-project/my-image:latest does not exist, or the service account used by the Workbench instance does not have storage.objectViewer access to the container registry.
B.The container registry endpoint is blocked by a firewall rule that does not allow egress to gcr.io.
C.The instance's underlying Compute Engine resources are exhausted, causing the container creation to timeout.
D.The Workbench instance is using an outdated custom image that is not compatible with the latest runtime version.
AnswerA

The container image gcr.io/my-project/my-image:latest does not exist, or the service account used by the Workbench instance does not have storage.objectViewer access to the container registry. This would prevent the instance from pulling the image, causing an UNHEALTHY status.

Why this answer

The UNHEALTHY status in Vertex AI Workbench typically occurs when the instance fails to start its container. Option A is correct because the most likely cause is that the specified container image (gcr.io/my-project/my-image:latest) does not exist in Container Registry, or the service account attached to the instance lacks the storage.objectViewer role on the registry bucket. Without this permission, the instance cannot pull the image, leading to a container creation failure and an UNHEALTHY state.

Options B, C, and D are less likely given the focus on the container image in the exhibit.

Exam trap

Google Cloud often tests the distinction between container image availability/permissions and network-level issues; the trap here is that candidates may assume a firewall or resource exhaustion is the cause, but the exhibit's focus on a specific container image points directly to a missing image or insufficient IAM permissions on the Container Registry.

How to eliminate wrong answers

Option B is wrong because while a firewall blocking egress to gcr.io could cause a pull failure, the exhibit does not mention any firewall rules, and the question asks for the 'most likely' cause based on the exhibit—lack of image existence or permissions is a more common and direct issue. Option C is wrong because Compute Engine resource exhaustion (e.g., CPU/memory) would typically cause a timeout or error during instance creation, not a persistent UNHEALTHY status after the instance is running; Vertex AI Workbench handles resource allocation separately. Option D is wrong because an outdated custom image would likely cause compatibility warnings or startup failures, but the exhibit shows a specific container image reference (gcr.io/my-project/my-image:latest), not a custom image issue; the UNHEALTHY status is tied to container pull failures, not image version mismatches.

55
MCQmedium

A company uses Cloud Logging to monitor their application logs. They notice that some logs from their Compute Engine instances are missing. The instances have the required logging permission. What is the most likely cause?

A.The log sink is not configured correctly.
B.The logging agent is not configured to send logs to Cloud Logging.
C.The instances are using a custom image without the logging agent.
D.The log bucket is in a different project.
E.The log entries are being filtered by the exclusion filter.
AnswerB

Compute Engine instances need the Cloud Logging agent installed and configured to forward logs; permissions alone do not transmit them. Without that agent configuration, logs remain local and never reach Cloud Logging, explaining the missing entries despite adequate IAM permissions.

Why this answer

Compute Engine instances do not automatically send logs to Cloud Logging. They require the Cloud Logging agent (based on fluentd) to be installed and configured to forward logs. Even with correct IAM permissions, without the agent, logs will not be collected.

Option B correctly identifies this missing agent as the most likely cause.

Exam trap

Google Cloud often tests the distinction between log collection (agent) and log routing (sinks) — the trap here is that candidates assume IAM permissions alone are sufficient, overlooking the mandatory agent installation and configuration step.

How to eliminate wrong answers

Option A is wrong because a log sink controls where logs are routed (e.g., to BigQuery or Pub/Sub), not whether logs are collected from instances; missing logs are a collection issue, not a routing issue. Option C is wrong because while a custom image might lack the agent, the question states the instances have the required logging permission, implying the agent could be installed separately; the most likely cause is the agent not being configured, not the image itself. Option D is wrong because log buckets in a different project would still receive logs if the sink is configured correctly; the issue is logs not appearing at all, not appearing in the wrong project.

Option E is wrong because exclusion filters remove logs after they are ingested; if logs are missing entirely, they were never ingested, so exclusion is not the cause.

56
MCQmedium

A media company runs a batch transcoding pipeline on Google Kubernetes Engine. Jobs read input from a Cloud Storage bucket and write output to a second bucket. The team wants the pipeline to keep processing through transient Cloud Storage 429 and 503 errors without losing work, and they want the pods to stop being killed mid-job during node upgrades. Which combination should the architect implement?

A.Increase the Cloud Storage bucket's requester pays setting and enable Object Versioning on both buckets.
B.Add exponential backoff retries with jitter in the application's Cloud Storage client, and configure a PodDisruptionBudget for the job pods.
C.Move the pipeline to a DaemonSet so one pod runs on every node and is not subject to eviction.
D.Set the pods' restartPolicy to Always and rely on the default pod eviction behavior during node upgrades.
AnswerB

Client-side retries with exponential backoff and jitter handle transient 429 and 503 responses from Cloud Storage without failing the unit of work, and a PodDisruptionBudget limits how many job pods can be voluntarily evicted at once during node upgrades. Together they address both the API flakiness and the mid-job termination risk during planned maintenance.

Why this answer

Reliability for this pipeline requires handling transient Cloud Storage errors at the client layer and protecting pods from simultaneous voluntary disruption. Exponential backoff with jitter retries 429 and 503 responses without overwhelming the service, while a PodDisruptionBudget constrains how many job pods are evicted during node upgrades, letting the pipeline drain gracefully instead of losing work.

Exam trap

The trap here is assuming pod restart policies or bucket-level durability features handle transient API errors, when retry logic must live in the client and eviction must be bounded by a PodDisruptionBudget.

57
Multi-Selecthard

A team is designing a disaster recovery (DR) plan for a critical application. Which THREE components are essential for a robust DR plan? (Choose 3)

Select 3 answers
A.Failover procedures and runbooks
B.Regular backups to a separate region
C.A single-region deployment for consistency
D.Monitoring and alerting for disaster events
E.Load testing to validate performance
AnswersA, B, D

Well-documented failover steps ensure quick recovery.

Why this answer

Failover procedures and runbooks (A) are essential because they provide step-by-step instructions for executing a controlled transition to the secondary site, ensuring minimal downtime and consistent recovery actions. Without documented runbooks, teams risk misconfigurations during a disaster, which can extend recovery time objectives (RTO) beyond acceptable limits.

Exam trap

Google Cloud often tests the misconception that a single-region deployment is acceptable for DR if it has high availability within that region, but the exam emphasizes that DR requires geographic separation to survive a full regional failure.

58
MCQeasy

A developer wants to monitor the CPU usage of a single Compute Engine VM and receive alerts when it exceeds 80%. What is the simplest way to achieve this?

A.Query the Compute Engine API periodically and check CPU usage.
B.Configure a Cloud Logging sink to BigQuery and set a scheduled query to detect high CPU.
C.Install the Cloud Monitoring agent and create an alerting policy based on the metric 'cpu.utilization'.
D.Use the managed instance group's autoscaling metric to trigger a notification.
AnswerC

The Monitoring agent collects CPU utilization from the OS and sends it to Cloud Monitoring, where you can set alerts.

Why this answer

The Cloud Monitoring agent (formerly Stackdriver agent) collects CPU utilization metrics from Compute Engine VMs and sends them to Cloud Monitoring. You can then create an alerting policy directly on the metric 'cpu.utilization' with a threshold of 80% without any custom scripting or additional infrastructure. This is the simplest and most native approach for a single VM.

Exam trap

Google Cloud often tests the misconception that you need to export logs to BigQuery or query APIs manually, when in fact the Cloud Monitoring agent provides a built-in, agent-based metric that can be alerted on directly.

How to eliminate wrong answers

Option A is wrong because periodically querying the Compute Engine API for CPU usage is inefficient, requires custom code, and does not provide real-time alerting; the API does not expose high-frequency CPU metrics natively. Option B is wrong because exporting logs to BigQuery and running scheduled queries adds unnecessary complexity, latency, and cost; Cloud Logging sinks are for log data, not for real-time metric-based alerting. Option D is wrong because managed instance group autoscaling metrics are designed for scaling groups of VMs, not for alerting on a single VM's CPU usage; they do not trigger notifications directly.

Ready to test yourself?

Try a timed practice session using only Ensure solution and operations reliability questions.