Courseiva

PCA · topic practice

Ensure solution and operations reliability practice questions

This domain covers keeping GKE, Compute Engine, and Cloud Run workloads available and observable. Expect scenario questions on Cloud Monitoring alerting policies, uptime checks, SLOs, multi-region failover, backup/restore, and capacity planning. You must choose the simplest reliable mechanism and justify trade-offs between cost, latency, and recovery objectives.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Ensure solution and operations reliability

What the exam tests

What to know about Ensure solution and operations reliability

Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.

Creating Cloud Monitoring alerting policies on CPU, latency, and error-rate metrics for VMs and GKE

Designing DR with multi-region deployments, Cloud Storage backups, and tested RTO/RPO targets

Configuring GKE liveness, readiness, and startup probes plus PodDisruptionBudgets and node auto-repair

Using Cloud Logging, Error Reporting, and uptime checks to detect and diagnose production incidents

Watch out for

Common Ensure solution and operations reliability exam traps

  • ▸Assuming zonal GKE clusters survive a zone outage; only regional clusters spread replicas across zones automatically.
  • ▸Setting alerts on raw metrics without understanding that sustained-duration conditions reduce false positives.
  • ▸Treating backups alone as a DR plan while ignoring failover runbooks, DNS cutover, and regular restore testing.

Practice set

Ensure solution and operations reliability questions

20 questions · select your answer, then reveal the explanation

A company is designing a disaster recovery plan for a Cloud SQL for PostgreSQL instance. They want to failover to a different region with minimal data loss and recovery time under 10 minutes. The database is 500 GB and experiences 2,000 write transactions per second. Which solution should they use?

A company deploys a stateful workload using StatefulSets on GKE. They want to ensure that if a pod is evicted, its persistent volume claim (PVC) is reattached to the replacement pod in the same zone. Which configuration achieves this?

A company runs a web application on Compute Engine behind an HTTP load balancer. They want to improve reliability by implementing failover across two regions. Which TWO actions should they take?

A company uses Cloud CDN to accelerate content delivery. They notice that some users receive stale content even after purging the cache. Which THREE factors could cause this?

A company deploys a critical application on Google Kubernetes Engine (GKE) and wants to ensure high availability during cluster upgrades. Which TWO practices should they follow?

A company runs a multi-tier application on Google Cloud: a frontend on App Engine Standard, a backend on Cloud Run, and a Cloud SQL database. The application experiences intermittent 500 errors when users submit forms. The errors correlate with high CPU usage on the Cloud SQL instance (db-n1-standard-2, 7.5 GB memory). The Cloud Run service has a concurrency setting of 80 and a maximum of 10 instances. The App Engine service uses automatic scaling. The team has verified that the application code is not the issue. They suspect the database is hitting connection limits. Current max_connections on Cloud SQL is 250. The Cloud Run service uses a connection pool of 10 connections per instance. The App Engine service uses a connection pool of 5 connections per instance. They also have a few batch jobs that run occasionally, using up to 10 connections. The team wants to resolve the errors with minimal cost and complexity. Which course of action should they take?

A company uses Cloud SQL for MySQL to host its production database. The database experiences high read traffic. The team wants to improve read performance without modifying the application. What should they do?

A company is running a critical application on Compute Engine. The application writes logs to a local persistent disk. The operations team wants to ensure logs are not lost if the VM fails. What should they do?

Which TWO options are best practices for ensuring high availability of an application running on Google Kubernetes Engine (GKE)?

A company runs a batch processing workload on Compute Engine that processes financial transactions. The workload runs daily and must complete within a 4-hour window. The application reads input data from Cloud Storage, processes it, and writes output to another Cloud Storage bucket. The current implementation uses a single VM with a 500 GB persistent disk. Recently, the data volume has increased, and the job is now taking over 6 hours, exceeding the SLA. The team is tasked with redesigning the solution to be faster and more reliable. They want to minimize costs and operational overhead. The data is critical and must not be lost. Which approach should they take?

Your company runs an e-commerce platform on Google Cloud. The application is deployed on Compute Engine instances in a managed instance group (MIG) with autoscaling based on CPU utilization. The database uses Cloud SQL for MySQL with a single instance. During a recent flash sale, traffic spiked and the application became slow, resulting in a poor user experience. After analyzing the incident, you discovered that the MIG scaled up but the Cloud SQL instance reached its maximum connections limit, causing some requests to fail. You need to recommend a solution to improve the reliability of the application for future traffic spikes. What should you do?

Which TWO actions should you take to improve the reliability of a stateful application deployed on Compute Engine with regional persistent disks?

Match each GCP networking concept to its definition.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Virtual Private Cloud for isolated network

Regional IP address range within a VPC

Controls ingress/egress traffic

Dynamically exchange routes using BGP

Connect two VPCs privately

An application uses Cloud Pub/Sub for asynchronous processing. Subscribers occasionally fail to acknowledge messages within the ack deadline, causing redelivery. How to improve reliability and prevent message buildup?

A global application uses Cloud Spanner with a multi-region configuration. During a regional outage, some transactions are failing. What is the recommended approach to maintain write availability?

A startup uses Cloud Functions for event-driven processing. They notice some functions are timing out. How to increase reliability without changing the business logic?

An organization wants to define an SLO for their API hosted on Cloud Endpoints. Which metric should they use as a Service Level Indicator (SLI) for availability?

A company uses Cloud Armor to protect their HTTP Load Balancer from DDoS attacks. During a traffic spike from a legitimate source, legitimate requests are being blocked. How should they tune the security policy to minimize false positives?

After a data corruption incident, a company needs to restore their Cloud SQL for PostgreSQL instance from a backup. What is the correct procedure to minimize downtime?

Question 20mediummultiple choice
Review the full routing breakdown →

A company is deploying a critical application on Compute Engine with an HTTP load balancer. They want to ensure that if an instance health check fails, traffic is automatically rerouted to healthy instances. Which configuration should they implement?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Ensure solution and operations reliability sessions

Start a Ensure solution and operations reliability only practice session

Every question in these sessions is drawn from the Ensure solution and operations reliability domain — nothing else.

Related practice questions

Related PCA topic practice pages

Move into related areas when this topic feels solid.

Analysing and Optimising Technical and Business Processes practice questions

Analysing and Optimising Technical and Business Processes practice questions for PCA.

Managing Implementation and Ensuring Solution and Operations Reliability practice questions

Targeted PCA practice covering Managing Implementation and Ensuring Solution and Operations Reliability.

Managing and Provisioning a Solution Infrastructure practice questions

Practise PCA questions linked to Managing and Provisioning a Solution Infrastructure.

Designing for Security and Compliance practice questions

Work through PCA questions on Designing for Security and Compliance.

Design for security and compliance practice questions

Design for security and compliance practice questions for PCA.

Design and plan a cloud solution architecture practice questions

Work through PCA questions on Design and plan a cloud solution architecture.

Manage and provision cloud infrastructure practice questions

Manage and provision cloud infrastructure practice questions for PCA.

Analyze and optimize technical and business processes practice questions

Analyze and optimize technical and business processes practice questions for PCA.

Ensure solution and operations reliability practice questions

Sharpen your PCA knowledge of Ensure solution and operations reliability.

Manage implementation of cloud architecture practice questions

Work through PCA questions on Manage implementation of cloud architecture.

PCA fundamentals practice questions

Practise PCA questions linked to PCA fundamentals.

PCA scenario practice questions

Work through PCA questions on PCA scenario.

Frequently asked questions

What does the PCA exam test about Ensure solution and operations reliability?
Be able to pick the right reliability control: Monitoring alerting policies for thresholds, regional GKE and multi-region services for availability, and tested backup/restore for DR. The single most important thing is matching every design to explicit RTO and RPO targets.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Ensure solution and operations reliability questions in a focused session?
Yes — the session launcher on this page draws every question from the Ensure solution and operations reliability domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PCA topics?
Use the topic links above to move to related areas, or go back to the PCA question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PCA exam covers. They are not copied from any real exam or dump site.