Courseiva

PCA · domain

Ensure solution and operations reliability

This domain covers keeping GKE, Compute Engine, and Cloud Run workloads available and observable. Expect scenario questions on Cloud Monitoring alerting policies, uptime checks, SLOs, multi-region failover, backup/restore, and capacity planning. You must choose the simplest reliable mechanism and justify trade-offs between cost, latency, and recovery objectives.

58 questions16 easy22 medium20 hard

Focused practice

Practice Ensure solution and operations reliability questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Ensure solution and operations reliability

Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.

Creating Cloud Monitoring alerting policies on CPU, latency, and error-rate metrics for VMs and GKE

Designing DR with multi-region deployments, Cloud Storage backups, and tested RTO/RPO targets

Configuring GKE liveness, readiness, and startup probes plus PodDisruptionBudgets and node auto-repair

Using Cloud Logging, Error Reporting, and uptime checks to detect and diagnose production incidents

Watch out for

Common Ensure solution and operations reliability exam traps

  • ▸Assuming zonal GKE clusters survive a zone outage; only regional clusters spread replicas across zones automatically.
  • ▸Setting alerts on raw metrics without understanding that sustained-duration conditions reduce false positives.
  • ▸Treating backups alone as a DR plan while ignoring failover runbooks, DNS cutover, and regular restore testing.

Question index

All Ensure solution and operations reliability questions (58)

Click any question to see the full explanation, or start a practice session above.

1

Your company's global e-commerce platform uses a managed instance group (MIG) in us-central1 and a Cloud Load Balancer. Traffic has grown, and you want to improve availability by distributing load across multiple regions. What should you do?

Medium
2

A company wants to monitor the health of their Cloud Run services. Which THREE metrics should they use to define a comprehensive health SLI? (Choose 3)

Easy
3

A company uses Cloud SQL for PostgreSQL. They want to minimize downtime during maintenance. Which feature should they enable?

Easy
4

Your company runs a critical application on Google Kubernetes Engine (GKE) with 5 nodes. The application experiences intermittent high latency every Friday afternoon. The team has ruled out infrastructure issues and suspects the application logic. You need to instrument the application to identify the root cause. Which approach should you take?

Easy
5

A financial services firm runs a latency-sensitive trading API on Google Kubernetes Engine (GKE). During peak market hours, the API occasionally returns errors because pods are evicted when nodes run out of memory. The team wants the workload to be protected from node-level resource pressure and to receive a graceful termination window when the node must be drained. Which configuration should they apply to the Deployment?

Hard
6

A company has deployed a critical application on Google Kubernetes Engine (GKE) with a Regional cluster (us-central1). The application uses a Cloud SQL for PostgreSQL database with a cross-region replica for disaster recovery. The SRE team needs to ensure that the application can survive a regional outage with minimal data loss. Which TWO actions should the team take to improve the reliability of the solution?

Medium
7

A company monitors their application with Cloud Monitoring. They set up an alerting policy to notify the on-call team when the 99th percentile latency exceeds 500 ms for 5 minutes. However, they receive false positive alerts due to short bursts. How should they refine the policy?

Medium
8

A company uses Cloud Storage for backup data. They want to protect against accidental deletion. Which option is best?

Easy
9

A company runs a global e-commerce site on GKE. They want to ensure disaster recovery with multi-region deployment. What is the best practice for configuring GKE clusters?

Easy
10

A logistics company runs a Cloud Run service that processes shipment events. They want to be notified and to trigger an automated rollback when the error rate of a new revision exceeds a threshold shortly after deployment. Which Google Cloud feature should they use?

Easy
11

Your company runs a stateless web application on Compute Engine. You want to ensure that if a zone fails, the application continues to serve traffic with minimal manual intervention. What should you do?

Easy
12

A company uses Cloud Storage to store user-uploaded content. They want to ensure that the data is highly durable and protected against accidental deletion. Which two features should they enable? (Choose two.)

Easy
13

You are deploying a new version of a microservice to Google Kubernetes Engine (GKE). The service must remain available during the rollout, and you need to minimize the risk of exposing bugs to all users at once. You want to gradually shift traffic to the new version while monitoring key metrics. Which strategy should you use?

Hard
14

Drag and drop the steps to set up a Cloud VPN tunnel between Google Cloud and an on-premises network into the correct order.

Medium
15

A company wants to monitor their Cloud Run services for errors and latency. Which Google Cloud product should they use?

Easy
16

Your service has a 99.99% uptime SLO (monthly error budget ~ 4 minutes). Which TWO monitoring practices best support this SLO? (Choose 2)

Hard
17

A company runs a stateful application on GKE using StatefulSets. Which THREE practices improve reliability?

Hard
18

You are running a Kubernetes cluster in GKE with the default node pool configuration shown in the exhibit. Your application requires high disk I/O performance. You notice that the application is experiencing high latency for disk operations. What is the most likely cause?

Hard
19

Your company runs a critical multi-tier application: a global HTTP(S) load balancer, multiple regional managed instance groups (MIGs) for the web tier, and Cloud Spanner for the data tier. You need to design for zone-level and region-level failures. What architecture ensures the highest availability?

Hard
20

A retail company uses a Cloud SQL for MySQL instance with a single zone. The database is critical for order processing, and the company wants to minimize downtime if the zone hosting the instance fails. They also want to ensure that the application can continue to write data during a zonal failure without manual intervention. What should they do?

Medium
21

A healthcare company runs a patient portal on Cloud Run services in the us-central1 region. The compliance team requires that the portal remain readable during a regional outage and that failover to a secondary region occur without changing the public hostname. The portal's data is stored in Cloud SQL for PostgreSQL. Which design should the architect recommend?

Hard
22

You are the lead cloud architect for a startup that runs a web application on Google Kubernetes Engine (GKE) with a standard (zonal) cluster. The application is deployed with 3 replicas of a stateless frontend service. During a recent incident, a zone outage caused all GKE nodes to become unavailable, leading to application downtime of 45 minutes. You need to redesign the cluster to tolerate a single zone failure with no more than 5 minutes of downtime. Your budget allows for at most a 20% increase in compute costs. Which approach should you take?

Easy
23

The exhibit shows a managed instance group configuration. What is the primary purpose of the 'autoHealingPolicies' section?

Hard
24

A company runs a stateful workload on Compute Engine with regional persistent disks (PD). They need to implement a disaster recovery (DR) plan with a Recovery Point Objective (RPO) of less than 1 hour and Recovery Time Objective (RTO) of less than 4 hours. Which THREE steps should they include in their DR plan? (Choose three.)

Medium
25

An organization uses Cloud Functions (2nd gen) for event-driven processing. They notice that some functions fail with 'memory limit exceeded' errors during peak load. The function processes messages from Pub/Sub and writes to Firestore. What should they do to improve reliability without sacrificing throughput?

Hard
26

A company uses Cloud Spanner for a global financial application. They experience increased latency and transaction aborts during peak hours. Which measure should they take first to improve reliability?

Easy
27

A developer wants to monitor a custom application metric from their application running on GKE. What should they use?

Easy
28

Your company runs a customer-facing API on Cloud Run with a concurrency setting of 80. The API calls a backend Cloud Function that performs a heavy computation (2–5 seconds). During peak hours, the API experiences increased latency and some requests time out after 60 seconds. Monitoring shows that the Cloud Run max instances is set to 100, and the Cloud Function max instances is set to 10. The timeout for Cloud Run is set to 300 seconds. The Cloud Function's timeout is set to 540 seconds. You need to reduce end-to-end latency and prevent timeouts while minimizing cost. Which action is most effective?

Medium
29

A developer ran the above command to create a health check for a backend service. Which of the following should they do to resolve the error?

Medium
30

Your company runs a data pipeline on Google Cloud using Cloud Dataflow for streaming processing from Pub/Sub to BigQuery. The pipeline writes to a BigQuery table partitioned by day. The data is used for real-time dashboards. Recently, a spike in traffic caused the Dataflow pipeline to fall behind, and the dashboard displayed stale data. You need to design the pipeline to handle traffic spikes without data loss or long delays. The pipeline must be cost-efficient and use defaults where possible. Which solution should you implement?

Hard
31

The exhibit shows the output of a 'gcloud compute instances describe' command for an instance. What is the most likely impact on reliability if the host machine needs maintenance?

Medium
32

A company runs a web application on Google Kubernetes Engine (GKE) with Cluster Autoscaler enabled. During a traffic spike, the application becomes slow and some requests timeout. The cluster has sufficient CPU and memory headroom. What is the most likely cause and solution?

Medium
33

A healthcare analytics company runs a stateless API on a regional managed instance group behind an external Application Load Balancer. The SRE team wants to improve reliability and reduce customer-visible errors during zonal and instance failures. (Choose two.)

Hard
34

Refer to the exhibit. An application running on a GCE instance (ID: 1234567890) is unable to connect to a database at 10.0.0.1:5432. The logs show repeated 'Connection refused' errors. What is the most likely cause?

Medium
35

A retail company runs a public-facing API on Cloud Run in the europe-west1 region. During a marketing campaign, traffic tripled within minutes. The service remained available, but some requests returned HTTP 503 errors. The team wants to reduce the chance of 503 errors during future traffic spikes while keeping the deployment simple. What should they do?

Easy
36

The exhibit shows a Cloud Storage bucket configuration. What does this configuration ensure?

Medium
37

Refer to the exhibit. A Cloud Deploy pipeline has a release with two targets: staging and prod. The staging rollout succeeded, but the prod rollout failed with 'MANIFEST_INVALID'. What is the most likely cause of the failure?

Hard
38

A company deploys a microservices application on Google Kubernetes Engine (GKE). Pods in one deployment are frequently OOMKilled. The team sets memory requests and limits, but pods still crash. What is the most likely remaining cause?

Medium
39

A financial services company is migrating a monolithic Java application to Google Kubernetes Engine (GKE) for improved scalability and reliability. The application serves real-time trading data and has strict latency requirements. Post-migration, the team observes frequent pod restarts due to OutOfMemory (OOM) errors, increased latency during peak trading hours, and occasional database connection timeouts. The current setup uses a single GKE cluster with a node pool of n1-standard-4 machines, a stateless application deployed as a Deployment with resource requests and limits set to 512 Mi memory and 1 CPU. The database is a Cloud SQL PostgreSQL instance with 2 vCPUs and 7.5 GB memory, and applications connect using a hardcoded connection string. The team wants to ensure reliable operation under load and during node maintenance events. Which course of action best addresses the reliability issues?

Hard
40

Your team manages a service with a 99.9% uptime SLO over a 30-day window. The error budget for this period is 43 minutes. In the first week, outages consumed 30 minutes of the budget. You are planning a new release. What should you do?

Medium
41

A company deploys a web application on Compute Engine behind an HTTP Load Balancer. They want to ensure only healthy instances receive traffic. What should they configure?

Easy
42

A company wants to improve the reliability of their microservices architecture on Google Cloud. Which TWO practices should they implement? (Choose 2)

Medium
43

A company runs a microservices-based application on Google Kubernetes Engine (GKE) with a Regional cluster. They want to improve reliability by implementing best practices for pod scheduling and resilience. Which TWO actions should they take? (Choose two.)

Hard
44

A company runs a stateful application on Compute Engine with persistent disks. They want to ensure data durability across a zone failure. What is the best approach?

Hard
45

Your organization is implementing a Disaster Recovery plan for a critical database. Which THREE components are essential for a robust DR strategy? (Choose 3)

Medium
46

An organization is migrating a legacy monolithic application to Google Cloud. The application currently runs on a single server with an on-premises database. The application is stateful and requires low-latency access to the database. The migration must minimize downtime and ensure high availability. Which architecture should the company adopt?

Hard
47

A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. During a traffic spike, some instances become unhealthy but are not automatically replaced. What is the most likely cause?

Medium
48

Which THREE options are valid strategies for disaster recovery (DR) in Google Cloud?

Hard
49

A company runs a critical application on Compute Engine instances in a managed instance group (MIG) with autoscaling. Users report intermittent 503 errors during traffic spikes. Which action should the company take to improve reliability?

Easy
50

You are responsible for incident management for a production service. You want to reduce manual toil during the initial response to common issues like high latency. What is the best approach?

Hard
51

Your organization uses Cloud Spanner for a customer database with a 99.999% availability SLA. You need a Disaster Recovery plan that ensures data consistency with zero RPO in case of a region failure. What should you do?

Medium
52

A retail company runs a customer-facing API on a managed instance group (MIG) of Compute Engine VMs behind an external Application Load Balancer. The SRE team wants the load balancer to stop sending traffic to a VM as soon as the local application health endpoint starts returning HTTP 500, even before the VM is fully unresponsive. Which load balancer component must be configured to achieve this?

Medium
53

You are responsible for a Cloud Run service that experiences occasional cold starts, causing increased latency. You want to minimize cold starts while keeping costs under control. What should you do?

Medium
54

You are investigating a Vertex AI Workbench instance (instance-2) that is showing UNHEALTHY status. Based on the exhibit, what is the most likely cause of the issue?

Hard
55

A company uses Cloud Logging to monitor their application logs. They notice that some logs from their Compute Engine instances are missing. The instances have the required logging permission. What is the most likely cause?

Medium
56

A media company runs a batch transcoding pipeline on Google Kubernetes Engine. Jobs read input from a Cloud Storage bucket and write output to a second bucket. The team wants the pipeline to keep processing through transient Cloud Storage 429 and 503 errors without losing work, and they want the pods to stop being killed mid-job during node upgrades. Which combination should the architect implement?

Medium
57

A team is designing a disaster recovery (DR) plan for a critical application. Which THREE components are essential for a robust DR plan? (Choose 3)

Hard
58

A developer wants to monitor the CPU usage of a single Compute Engine VM and receive alerts when it exceeds 80%. What is the simplest way to achieve this?

Easy

Frequently asked questions

What does the Ensure solution and operations reliability domain cover on the PCA exam?
Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.
How many questions are in this domain?
This page lists all 58 Ensure solution and operations reliability questions in the PCA question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Ensure solution and operations reliability questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pca GOOGLE-PCA reliability ops Practice Questions