PCA · domain
Managing Implementation and Ensuring Solution and Operations Reliability
This domain covers deploying, monitoring, and operating workloads on Google Cloud: Cloud Monitoring alerting policies and notification channels, Cloud SQL high availability and failover, GKE release and rollback strategy, and migration tooling. Questions are scenario-based, asking you to pick the correct service configuration, migration approach, or change-management practice for a stated reliability or downtime requirement.
Focused practice
Practice Managing Implementation and Ensuring Solution and Operations Reliability questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Managing Implementation and Ensuring Solution and Operations Reliability
You must be able to configure Cloud Monitoring alerts with several notification channels, enable Cloud SQL high availability for cross-zone failover, and design staged GKE rollouts with rollback. The key is matching the reliability or downtime requirement to the specific feature that actually delivers it.
Cloud Monitoring alerting policies with multiple notification channels including email and Pub/Sub
Cloud SQL high availability configuration with a standby in a different zone for automatic failover
GKE rollout strategies: canary or blue/green deployments with quick rollback and pre-rollout validation
Database Migration Service for phased, low-downtime, consistent migration of on-premises databases to Cloud SQL
Watch out for
Common Managing Implementation and Ensuring Solution and Operations Reliability exam traps
- ▸Assuming a single alerting policy can only notify one channel; Cloud Monitoring policies support multiple notification channels simultaneously.
- ▸Enabling backups or read replicas and expecting automatic failover; only high availability configuration provides a standby and failover.
- ▸Rolling out GKE changes directly to all production pods without a staged canary and a tested rollback path.
Question index
All Managing Implementation and Ensuring Solution and Operations Reliability questions (78)
Click any question to see the full explanation, or start a practice session above.
An engineer needs to view the logs of a specific Compute Engine instance in near real-time from the command line. Which gcloud command should they use?
Easy2An organization wants to connect their on-premises data center to Google Cloud with a dedicated 10 Gbps link. They require high availability and have budget for two physically diverse connections. Which solution should they choose?
Easy3A media company runs a video transcoding service on GKE Standard. The service experiences sudden traffic spikes, and the operations team wants to ensure that the cluster can scale nodes automatically and that pods are rescheduled quickly when a node fails. The team also wants to monitor and alert on resource saturation. Which two actions should the cloud architect take to meet these requirements? (Choose two.)
Medium4You need to monitor the performance of a production Cloud Run service and set an alert when the p99 latency exceeds 500 ms over a 5-minute window. Which combination of Cloud Monitoring resources should you use?
Medium5A small startup is deploying a new application on Google Cloud. They want to ensure that they can monitor the application's performance and receive alerts when certain thresholds are exceeded. They have limited operational staff and want a managed solution that requires minimal configuration. Which Google Cloud service should they use?
Easy6Your company runs a stateful application on Compute Engine instances in a managed instance group (MIG). The application writes data to a persistent disk attached to each instance. You need to ensure that the application can automatically recover from a zone failure by recreating instances in another zone with their persistent disks. You also want to minimize data loss. Which configuration should you implement?
Medium7An organization is using Cloud Interconnect to connect their on-premises network to Google Cloud. They need to ensure 99.99% availability for their connection. Which configuration meets this requirement?
Medium8Your team runs a stateful analytics workload on a Managed Instance Group (MIG) of Compute Engine VMs. The VMs write intermediate results to local SSD scratch disks. During a recent incident, an autoscaling event terminated VMs and the intermediate data was lost, causing hours of recomputation. You need to change the deployment so that when a VM is terminated by the autoscaler, a shutdown script has enough time to flush the intermediate results to a Cloud Storage bucket before the VM is deleted. What should you do?
Medium9During a load test, an application running on GKE experiences high latency and errors. You suspect the issue is due to insufficient cluster resources. Which gcloud command should you use to quickly check the current resource utilization of all nodes in the cluster?
Hard10You are responsible for ensuring the reliability of a high-traffic web application running on Google Kubernetes Engine (GKE). You need to implement a monitoring strategy that alerts you when the application's error rate exceeds 1% over a 5-minute window. You want to minimize false positives and ensure alerts are actionable. What should you do?
Hard11To achieve a 99.999% availability SLA for a globally distributed application using Cloud Spanner, which configuration is required?
Easy12An organization needs to implement a change management process for a mission-critical application on GKE. They want to validate performance before full rollout and be able to roll back quickly. Which THREE practices should they adopt? (Choose THREE.)
Medium13A company is planning a phased migration of their on-premises database to Cloud SQL. They want to minimize downtime and ensure data consistency. Which approach should they use?
Medium14An application running on Compute Engine is experiencing increased latency. You suspect a network bottleneck due to high egress traffic. Which gcloud command can you use to quickly check the network egress traffic for a specific VM instance?
Hard15Your organization operates a multi-project Google Cloud environment. A security team requires that any new Compute Engine instance created in the production folder must have OS Login enabled and must not use project-wide SSH keys. You want to enforce this centrally with the least operational overhead and ensure that non-compliant creation attempts are denied. What should you do?
Hard16Your team is deploying a new version of a microservices application on Google Kubernetes Engine (GKE). You want to gradually shift traffic to the new version while monitoring key performance indicators (KPIs) such as error rate and latency. If KPIs degrade, you need to automatically roll back. Which approach should you use?
Medium17You are responsible for operations reliability of a production service running on Google Cloud. The service is deployed on GKE and exposes an external HTTPS endpoint through an external Application Load Balancer. You need to implement monitoring that detects when the service is unhealthy from the user's perspective and alerts the on-call team. (Choose two.)
Medium18A startup runs a customer-facing web application on Cloud Run. The operations team needs to know when the service's request latency exceeds a threshold so they can respond before users complain. They want to be notified by email and also want a record of the incident for later review. Which Google Cloud service should they use to define the alerting policy?
Easy19Your company runs a multi-region Cloud Spanner instance for a global financial application. The SLA requirement is 99.999% availability. You need to ensure that the database remains available during a regional outage. What configuration should you use?
Medium20Your company has a service running on Google Kubernetes Engine (GKE) that experiences occasional spikes in traffic. You need to ensure that the service remains available during these spikes by automatically scaling the number of pods based on CPU utilization. You also want to minimize cost by scaling down when traffic decreases. Which Kubernetes resource should you configure?
Easy21Your team is responsible for a production service running on Google Cloud. You need to define Service Level Objectives (SLOs) and monitor them using Cloud Monitoring. You want to be alerted when the service's error budget is being consumed too quickly. Which approach should you take?
Medium22A company is designing a disaster recovery (DR) plan for their Cloud SQL for PostgreSQL instance. They need to recover the database to a specific point in time within the last 7 days, with a Recovery Point Objective (RPO) of less than 1 hour. Which feature should they use?
Medium23Your company runs a production microservices application on GKE Standard. The operations team wants to be notified when any pod in the cluster is repeatedly restarting, indicating a potential CrashLoopBackOff. They want to use Cloud Monitoring to create an alert that fires when a container restarts more than 5 times in a 10-minute window. Which metric should they use as the basis for the alerting policy?
Medium24A retail company runs a web application on Compute Engine instances behind an external HTTP(S) load balancer. During a flash sale, the operations team notices that the load balancer is returning HTTP 502 errors for a subset of requests. The backend service health checks are passing, and the instances are not under heavy CPU load. The team wants to identify the root cause quickly and prevent recurrence. Which action should they take first?
Medium25Your company is deploying a new application on Google Cloud and needs to ensure that it can meet a 99.9% availability SLA. You are designing the architecture for high availability. Which two practices should you implement? (Choose two.)
Medium26Your team manages a production web application on Compute Engine behind an external Application Load Balancer. During a recent incident, the load balancer's backend service marked all instances as unhealthy because the health check endpoint returned HTTP 200 but the application was actually in a degraded state. You need Cloud Monitoring to alert the operations team when the application's error rate exceeds 5% over a 5-minute window. You also need to ensure that the alert does not fire during planned maintenance windows. Which approach should you take?
Medium27Your company runs a production application on Compute Engine instances behind a managed instance group (MIG). You need to perform a rolling update with canary testing, gradually shifting traffic to the new version only if performance metrics are healthy. Which approach should you use?
Hard28A company is designing a highly available architecture for a web application using Google Cloud. They need to ensure that the application remains available even if an entire Google Cloud region experiences an outage. Which THREE components should they include in their architecture? (Choose THREE.)
Hard29A company wants to connect their on-premises network to Google Cloud with a 99.99% SLA using encrypted tunnels over the public internet. Which connectivity solution should they choose?
Easy30A startup deploys a containerized web application on Cloud Run. They want to release a new revision to a small percentage of users before promoting it to all traffic, and they need the ability to roll back instantly if errors increase. Which Cloud Run feature should they use?
Easy31Your company has a Service Level Objective (SLO) of 99.9% availability for a web application running on Google Cloud. You want to create an alert that notifies the on-call team when the error budget is being consumed too quickly. Which Google Cloud service should you use?
Easy32A company wants to set up monitoring and alerting for their application running on GKE. They need to receive alerts via email and also trigger an automated remediation workflow. Which TWO components should they use? (Choose two.)
Medium33A financial services company runs a payment API on Compute Engine behind an internal passthrough Network Load Balancer. The compliance team requires that all administrative actions on the project be attributable to a named human, that production changes be reviewed before taking effect, and that no single engineer can delete the production database. Which combination of Google Cloud controls should the cloud architect implement?
Hard34You are designing a disaster recovery plan for a critical application running on GKE. You need to back up the cluster's state and application data. Which TWO services should you use together? (Choose 2)
Medium35Your company plans to connect an on-premises data center to Google Cloud with a Dedicated Interconnect. You need to ensure high availability for the connection. What is the minimum configuration required to meet a 99.99% SLA for Dedicated Interconnect?
Medium36A company wants to connect their on-premises data center to Google Cloud with a dedicated private connection that provides 99.99% availability and supports up to 100 Gbps bandwidth. They have a colocation facility near a Google Cloud region. Which connectivity option should they choose?
Medium37A healthcare company runs a critical patient portal on Google Kubernetes Engine. The security team requires that all container images be scanned for vulnerabilities before deployment, that only images from a trusted registry be admitted to the cluster, and that any attempt to deploy an untrusted image be blocked and logged. Which Google Cloud feature should the cloud architect implement to enforce these admission requirements?
Hard38A company uses Cloud Logging to capture application logs. They need to alert when the number of errors exceeds 100 in a 5-minute window. Which type of alert should they create?
Easy39A healthcare company runs a critical application on Google Kubernetes Engine (GKE) that processes patient data. The compliance team requires that all container images be scanned for vulnerabilities before deployment, and that only images from a trusted registry be allowed. The security team wants to enforce this policy across all clusters in the organization. They also need to audit any attempts to deploy untrusted images. Which combination of Google Cloud services should they use?
Hard40An application running on GKE Autopilot is experiencing intermittent failures due to resource limits. The team wants to ensure that the application always has enough CPU and memory without manual node management. What should they do?
Hard41An engineer needs to create a custom dashboard in Cloud Monitoring to track the 99th percentile latency of their application over the last 7 days. Which type of metric should they use?
Easy42A financial services company runs a critical PostgreSQL database on Cloud SQL. They need to ensure automatic failover to a replica in another zone within the same region with minimal data loss. What configuration should they choose?
Hard43A company uses Cloud SQL for MySQL for its transactional database. They need to ensure automatic failover in case of a zonal outage with minimal data loss. What configuration should they use?
Medium44A company needs to retain object versions in Cloud Storage for 90 days to protect against accidental deletion or modification. After 90 days, versions should be deleted. What feature should they enable?
Easy45A financial services company runs a latency-sensitive trading application on GKE. The platform team must guarantee that the application can be recovered within a 15-minute recovery time objective (RTO) and a 5-minute recovery point objective (RPO) after a regional failure. They use a multi-region Cloud Storage bucket for configuration and a regional GKE cluster. Which additional design element is required to meet both objectives?
Hard46A company has a Cloud SQL for PostgreSQL instance in a single zone. To achieve high availability, they want to ensure automatic failover with zero data loss and minimal downtime. Which configuration should they use?
Medium47An e-commerce platform uses Cloud Spanner in a multi-region configuration. They want to achieve the highest possible availability SLA. Which deployment configuration should they choose?
Hard48Your team uses Cloud SQL for PostgreSQL for an e-commerce application. You want to perform point-in-time recovery (PITR) to recover from a logical error that occurred 10 minutes ago. Which prerequisites are required?
Medium49You want to create a log-based alert in Cloud Logging that triggers when a specific error message appears in application logs. What is the first step?
Easy50Your organization runs a global e-commerce platform on Google Kubernetes Engine (GKE). The security team requires that all container images deployed to the cluster are scanned for vulnerabilities and that deployments are blocked if critical vulnerabilities are found. They also want to minimize operational overhead. What should you do?
Medium51An engineer needs to troubleshoot a production issue on a Compute Engine instance. They suspect the instance is running out of memory. Which THREE actions should they take to diagnose the problem? (Choose THREE.)
Easy52An engineer needs to list all Compute Engine instances in a project using the command line. Which gcloud command should they use?
Easy53A team is adopting a DevOps model and wants to reduce the risk of configuration drift between environments. They deploy the same application to development, staging, and production projects on Google Cloud. Which practice should they adopt to ensure consistent, repeatable deployments across all environments?
Easy54Your organization runs a production Cloud SQL for PostgreSQL instance. You need to ensure that if the primary zone fails, the database automatically fails over to a standby with no data loss. Which configuration should you use?
Medium55Your company runs a stateful application on GKE that stores data in persistent volumes backed by Compute Engine persistent disks. You need to back up the application data and the Kubernetes resource configurations (deployments, services, etc.) for disaster recovery. Which tool should you use?
Medium56A company wants to implement a disaster recovery (DR) strategy for their Cloud SQL for MySQL databases. They need to be able to recover to a specific point in time (within seconds) in case of accidental data deletion. Which TWO actions should they take? (Choose TWO.)
Medium57Your company uses Cloud Spanner in a multi-region configuration to achieve 99.999% availability. You need to understand the impact of a regional failure on read and write availability. Which statement is correct?
Hard58A financial services company runs a latency-sensitive trading application on Compute Engine. The operations team needs to detect performance regressions and correlate them with recent deployments without instrumenting application code. They want to use Cloud Monitoring and Cloud Logging features that work automatically for Compute Engine VMs. (Choose two.)
Hard59You need to automatically roll back a GKE deployment if a new version causes a spike in 5xx errors. The deployment uses a canary strategy with Istio traffic splitting. What should you do?
Easy60A retail company runs an e-commerce platform on GKE. The SRE team wants to measure the error budget for a service level objective (SLO) defined as the proportion of requests served with HTTP 2xx or 3xx status over a 28-day window. They need a monitoring configuration that computes the burn rate and alerts when the budget is being consumed too quickly, while avoiding noisy alerts during brief spikes. What should they do?
Medium61Your team is following an incident management process. After resolving a major incident, you are tasked with conducting a postmortem. What is the PRIMARY goal of the postmortem process in Google Cloud's recommended approach?
Hard62You need to create a Cloud Logging sink that exports logs to a BigQuery dataset for long-term analysis. Which destination type should you specify?
Easy63Your organization wants to use Cloud SQL for a MySQL database with automatic failover in the event of a zone outage. Which configuration should you choose?
Easy64An organization needs to run a stateful application on Google Kubernetes Engine (GKE) where the nodes are fully managed by Google and the application workload SLAs are guaranteed. They want to minimize operational overhead. Which GKE mode should they use?
Medium65A team is migrating a monolithic application to microservices on GKE. They want to gradually shift users to the new microservices version while keeping the old monolithic version running. They need to route a small percentage of users based on a cookie. Which traffic management approach should they use?
Medium66An e-commerce company runs its order-processing service on Cloud Run. During flash sales, the service experiences sudden traffic spikes, and the operations team observes that new instances take too long to start, causing elevated latency and some request failures. The service has a large container image and initializes database connection pools at startup. Which configuration change should the team make to reduce cold-start impact while controlling cost?
Hard67Your organization requires that all production changes to Google Cloud resources be auditable and that you can identify who made a change and when. You need to configure logging to meet this requirement. What should you do?
Easy68Your company uses Cloud VPN (HA VPN) to connect to Google Cloud. You need to achieve a 99.99% SLA for the VPN connection. What configuration is required?
Medium69Your organization runs a critical application on Google Cloud that uses Cloud SQL for PostgreSQL. The database is in us-central1. The business requires a recovery point objective (RPO) of 5 minutes and a recovery time objective (RTO) of 1 hour in case of a regional failure. What should you do?
Hard70An organization wants to receive alerts when their Cloud SQL instance's CPU utilization exceeds 80% for 5 minutes. They want to send the alert to both email and a Pub/Sub topic for further processing. What should they do?
Medium71A company has a Cloud SQL for MySQL instance with automated backups enabled. They need to recover the database to a specific point in time within the last hour. Which feature should they use?
Medium72An organization uses Cloud Storage to store critical documents. They want to protect against accidental deletion or overwriting of objects. Which feature should they enable?
Medium73Your company has a production Cloud SQL for PostgreSQL instance in us-central1 with automated backups enabled. You need to ensure that if the zone fails, the database automatically fails over to a standby in a different zone with minimal downtime. What should you do?
Medium74Your team operates a production e-commerce application on a managed instance group (MIG) that serves traffic through a global external Application Load Balancer. During a new release, the team wants to deploy the new version to a small subset of instances and then progressively increase traffic to it while monitoring error rates, with the ability to immediately roll back if errors spike. The new version is already built as a custom image. Which approach should you use?
Medium75Your company wants to implement a canary deployment for a microservice running on GKE. You need to gradually shift traffic from the stable version to the canary version while monitoring error rates. Which THREE components or practices should you use? (Choose 3)
Hard76Your organization runs a microservices application on Google Kubernetes Engine (GKE). You need to ensure that the application can be rolled back quickly if a new deployment causes errors. You want to use a deployment strategy that allows you to shift traffic back to the previous version with minimal downtime. Which approach should you use?
Medium77Your organization stores critical financial data in Cloud Storage. You need to ensure that if an object is deleted or overwritten, you can recover it within 30 days. What feature should you enable?
Medium78A team needs to set up alerting for a production service. They want to receive notifications when the 99th percentile latency exceeds 500ms for 5 minutes. Which two Cloud Monitoring components are required? (Choose two.)
MediumOther domains
All PCA exam domains
Frequently asked questions
- What does the Managing Implementation and Ensuring Solution and Operations Reliability domain cover on the PCA exam?
- You must be able to configure Cloud Monitoring alerts with several notification channels, enable Cloud SQL high availability for cross-zone failover, and design staged GKE rollouts with rollback. The key is matching the reliability or downtime requirement to the specific feature that actually delivers it.
- How many questions are in this domain?
- This page lists all 78 Managing Implementation and Ensuring Solution and Operations Reliability questions in the PCA question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Managing Implementation and Ensuring Solution and Operations Reliability questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.