Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.
Start practicing
Ensure solution and operations reliability — choose a session length
Free · No account required
Domain overview
This domain covers keeping GKE, Compute Engine, and Cloud Run workloads available and observable. Expect scenario questions on Cloud Monitoring alerting policies, uptime checks, SLOs, multi-region failover, backup/restore, and capacity planning. You must choose the simplest reliable mechanism and justify trade-offs between cost, latency, and recovery objectives.
Exam objectives
Creating Cloud Monitoring alerting policies on CPU, latency, and error-rate metrics for VMs and GKE
Designing DR with multi-region deployments, Cloud Storage backups, and tested RTO/RPO targets
Configuring GKE liveness, readiness, and startup probes plus PodDisruptionBudgets and node auto-repair
Using Cloud Logging, Error Reporting, and uptime checks to detect and diagnose production incidents
Assuming zonal GKE clusters survive a zone outage; only regional clusters spread replicas across zones automatically.
Setting alerts on raw metrics without understanding that sustained-duration conditions reduce false positives.
Treating backups alone as a DR plan while ignoring failover runbooks, DNS cutover, and regular restore testing.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. During a traffic spike, some instances become unhealthy but are not automatically replaced. What is the most likely cause?
2A company uses Cloud Spanner for a global financial application. They experience increased latency and transaction aborts during peak hours. Which measure should they take first to improve reliability?
3A company deploys a microservices application on Google Kubernetes Engine (GKE). Pods in one deployment are frequently OOMKilled. The team sets memory requests and limits, but pods still crash. What is the most likely remaining cause?
4An organization uses Cloud Functions (2nd gen) for event-driven processing. They notice that some functions fail with 'memory limit exceeded' errors during peak load. The function processes messages from Pub/Sub and writes to Firestore. What should they do to improve reliability without sacrificing throughput?
5A company monitors their application with Cloud Monitoring. They set up an alerting policy to notify the on-call team when the 99th percentile latency exceeds 500 ms for 5 minutes. However, they receive false positive alerts due to short bursts. How should they refine the policy?
6A company runs a web application on Google Kubernetes Engine (GKE) with Cluster Autoscaler enabled. During a traffic spike, the application becomes slow and some requests timeout. The cluster has sufficient CPU and memory headroom. What is the most likely cause and solution?
7An organization is migrating a legacy monolithic application to Google Cloud. The application currently runs on a single server with an on-premises database. The application is stateful and requires low-latency access to the database. The migration must minimize downtime and ensure high availability. Which architecture should the company adopt?
8Which THREE options are valid strategies for disaster recovery (DR) in Google Cloud?
9A company has deployed a critical application on Google Kubernetes Engine (GKE) with a Regional cluster (us-central1). The application uses a Cloud SQL for PostgreSQL database with a cross-region replica for disaster recovery. The SRE team needs to ensure that the application can survive a regional outage with minimal data loss. Which TWO actions should the team take to improve the reliability of the solution?
10You are investigating a Vertex AI Workbench instance (instance-2) that is showing UNHEALTHY status. Based on the exhibit, what is the most likely cause of the issue?
11You are running a Kubernetes cluster in GKE with the default node pool configuration shown in the exhibit. Your application requires high disk I/O performance. You notice that the application is experiencing high latency for disk operations. What is the most likely cause?
12Your company runs a critical application on Google Kubernetes Engine (GKE) with 5 nodes. The application experiences intermittent high latency every Friday afternoon. The team has ruled out infrastructure issues and suspects the application logic. You need to instrument the application to identify the root cause. Which approach should you take?
13Drag and drop the steps to set up a Cloud VPN tunnel between Google Cloud and an on-premises network into the correct order.
14A company deploys a web application on Compute Engine behind an HTTP Load Balancer. They want to ensure only healthy instances receive traffic. What should they configure?
15A developer wants to monitor a custom application metric from their application running on GKE. What should they use?
16A company runs a stateful application on Compute Engine with persistent disks. They want to ensure data durability across a zone failure. What is the best approach?
17A company wants to improve the reliability of their microservices architecture on Google Cloud. Which TWO practices should they implement? (Choose 2)
18A team is designing a disaster recovery (DR) plan for a critical application. Which THREE components are essential for a robust DR plan? (Choose 3)
19A company wants to monitor the health of their Cloud Run services. Which THREE metrics should they use to define a comprehensive health SLI? (Choose 3)
20A company uses Cloud Logging to monitor their application logs. They notice that some logs from their Compute Engine instances are missing. The instances have the required logging permission. What is the most likely cause?
21A company uses Cloud Storage to store user-uploaded content. They want to ensure that the data is highly durable and protected against accidental deletion. Which two features should they enable? (Choose two.)
22A developer ran the above command to create a health check for a backend service. Which of the following should they do to resolve the error?
23Your company runs a stateless web application on Compute Engine. You want to ensure that if a zone fails, the application continues to serve traffic with minimal manual intervention. What should you do?
24A developer wants to monitor the CPU usage of a single Compute Engine VM and receive alerts when it exceeds 80%. What is the simplest way to achieve this?
25Your company's global e-commerce platform uses a managed instance group (MIG) in us-central1 and a Cloud Load Balancer. Traffic has grown, and you want to improve availability by distributing load across multiple regions. What should you do?
26Your organization uses Cloud Spanner for a customer database with a 99.999% availability SLA. You need a Disaster Recovery plan that ensures data consistency with zero RPO in case of a region failure. What should you do?
27Your team manages a service with a 99.9% uptime SLO over a 30-day window. The error budget for this period is 43 minutes. In the first week, outages consumed 30 minutes of the budget. You are planning a new release. What should you do?
28Your company runs a critical multi-tier application: a global HTTP(S) load balancer, multiple regional managed instance groups (MIGs) for the web tier, and Cloud Spanner for the data tier. You need to design for zone-level and region-level failures. What architecture ensures the highest availability?
29You are responsible for incident management for a production service. You want to reduce manual toil during the initial response to common issues like high latency. What is the best approach?
30Your organization is implementing a Disaster Recovery plan for a critical database. Which THREE components are essential for a robust DR strategy? (Choose 3)
31Your service has a 99.99% uptime SLO (monthly error budget ~ 4 minutes). Which TWO monitoring practices best support this SLO? (Choose 2)
32The exhibit shows the output of a 'gcloud compute instances describe' command for an instance. What is the most likely impact on reliability if the host machine needs maintenance?
33The exhibit shows a Cloud Storage bucket configuration. What does this configuration ensure?
34The exhibit shows a managed instance group configuration. What is the primary purpose of the 'autoHealingPolicies' section?
35A company runs a global e-commerce site on GKE. They want to ensure disaster recovery with multi-region deployment. What is the best practice for configuring GKE clusters?
36A company uses Cloud SQL for PostgreSQL. They want to minimize downtime during maintenance. Which feature should they enable?
37A company wants to monitor their Cloud Run services for errors and latency. Which Google Cloud product should they use?
38A company uses Cloud Storage for backup data. They want to protect against accidental deletion. Which option is best?
39A company runs a stateful application on GKE using StatefulSets. Which THREE practices improve reliability?
40A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. Users report intermittent 503 errors during traffic spikes. Which action should the company take to improve reliability?
41A company runs a microservices-based application on Google Kubernetes Engine (GKE) with a Regional cluster. They want to improve reliability by implementing best practices for pod scheduling and resilience. Which TWO actions should they take? (Choose two.)
42A company runs a stateful workload on Compute Engine with regional persistent disks (PD). They need to implement a disaster recovery (DR) plan with a Recovery Point Objective (RPO) of less than 1 hour and Recovery Time Objective (RTO) of less than 4 hours. Which THREE steps should they include in their DR plan? (Choose three.)
43You are the lead cloud architect for a startup that runs a web application on Google Kubernetes Engine (GKE) with a standard (zonal) cluster. The application is deployed with 3 replicas of a stateless frontend service. During a recent incident, a zone outage caused all GKE nodes to become unavailable, leading to application downtime of 45 minutes. You need to redesign the cluster to tolerate a single zone failure with no more than 5 minutes of downtime. Your budget allows for at most a 20% increase in compute costs. Which approach should you take?
44Your company runs a customer-facing API on Cloud Run with a concurrency setting of 80. The API calls a backend Cloud Function that performs a heavy computation (2–5 seconds). During peak hours, the API experiences increased latency and some requests time out after 60 seconds. Monitoring shows that the Cloud Run max instances is set to 100, and the Cloud Function max instances is set to 10. The timeout for Cloud Run is set to 300 seconds. The Cloud Function's timeout is set to 540 seconds. You need to reduce end-to-end latency and prevent timeouts while minimizing cost. Which action is most effective?
45Your company runs a data pipeline on Google Cloud using Cloud Dataflow for streaming processing from Pub/Sub to BigQuery. The pipeline writes to a BigQuery table partitioned by day. The data is used for real-time dashboards. Recently, a spike in traffic caused the Dataflow pipeline to fall behind, and the dashboard displayed stale data. You need to design the pipeline to handle traffic spikes without data loss or long delays. The pipeline must be cost-efficient and use defaults where possible. Which solution should you implement?
46A financial services company is migrating a monolithic Java application to Google Kubernetes Engine (GKE) for improved scalability and reliability. The application serves real-time trading data and has strict latency requirements. Post-migration, the team observes frequent pod restarts due to OutOfMemory (OOM) errors, increased latency during peak trading hours, and occasional database connection timeouts. The current setup uses a single GKE cluster with a node pool of n1-standard-4 machines, a stateless application deployed as a Deployment with resource requests and limits set to 512 Mi memory and 1 CPU. The database is a Cloud SQL PostgreSQL instance with 2 vCPUs and 7.5 GB memory, and applications connect using a hardcoded connection string. The team wants to ensure reliable operation under load and during node maintenance events. Which course of action best addresses the reliability issues?
47Refer to the exhibit. An application running on a GCE instance (ID: 1234567890) is unable to connect to a database at 10.0.0.1:5432. The logs show repeated 'Connection refused' errors. What is the most likely cause?
48Refer to the exhibit. A Cloud Deploy pipeline has a release with two targets: staging and prod. The staging rollout succeeded, but the prod rollout failed with 'MANIFEST_INVALID'. What is the most likely cause of the failure?
49A retail company uses a Cloud SQL for MySQL instance with a single zone. The database is critical for order processing, and the company wants to minimize downtime if the zone hosting the instance fails. They also want to ensure that the application can continue to write data during a zonal failure without manual intervention. What should they do?
50You are deploying a new version of a microservice to Google Kubernetes Engine (GKE). The service must remain available during the rollout, and you need to minimize the risk of exposing bugs to all users at once. You want to gradually shift traffic to the new version while monitoring key metrics. Which strategy should you use?
51A retail company runs a public-facing API on Cloud Run in the europe-west1 region. During a marketing campaign, traffic tripled within minutes. The service remained available, but some requests returned HTTP 503 errors. The team wants to reduce the chance of 503 errors during future traffic spikes while keeping the deployment simple. What should they do?
52A retail company runs a customer-facing API on a managed instance group (MIG) of Compute Engine VMs behind an external Application Load Balancer. The SRE team wants the load balancer to stop sending traffic to a VM as soon as the local application health endpoint starts returning HTTP 500, even before the VM is fully unresponsive. Which load balancer component must be configured to achieve this?
53You are responsible for a Cloud Run service that experiences occasional cold starts, causing increased latency. You want to minimize cold starts while keeping costs under control. What should you do?
54A financial services firm runs a latency-sensitive trading API on Google Kubernetes Engine (GKE). During peak market hours, the API occasionally returns errors because pods are evicted when nodes run out of memory. The team wants the workload to be protected from node-level resource pressure and to receive a graceful termination window when the node must be drained. Which configuration should they apply to the Deployment?
55A media company runs a batch transcoding pipeline on Google Kubernetes Engine. Jobs read input from a Cloud Storage bucket and write output to a second bucket. The team wants the pipeline to keep processing through transient Cloud Storage 429 and 503 errors without losing work, and they want the pods to stop being killed mid-job during node upgrades. Which combination should the architect implement?
56A healthcare company runs a patient portal on Cloud Run services in the us-central1 region. The compliance team requires that the portal remain readable during a regional outage and that failover to a secondary region occur without changing the public hostname. The portal's data is stored in Cloud SQL for PostgreSQL. Which design should the architect recommend?
57A healthcare analytics company runs a stateless API on a regional managed instance group behind an external Application Load Balancer. The SRE team wants to improve reliability and reduce customer-visible errors during zonal and instance failures. (Choose two.)
58A logistics company runs a Cloud Run service that processes shipment events. They want to be notified and to trigger an automated rollback when the error rate of a new revision exceeds a threshold shortly after deployment. Which Google Cloud feature should they use?
Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.
The Courseiva PCA question bank contains 58 questions in the Ensure solution and operations reliability domain, covering the 6% of the exam attributed to this domain in the official Google Cloud blueprint. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Ensure solution and operations reliability domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included