Courseiva

CCNA Managing Implementation and Ensuring Solution and Operations Reliability Questions

75 of 78 questions · Page 1/2 · Managing Implementation and Ensuring Solution and Operations Reliability · Answers revealed

1
MCQeasy

An engineer needs to view the logs of a specific Compute Engine instance in near real-time from the command line. Which gcloud command should they use?

Answer options not yet available.

Why this answer

gcloud logging tail streams logs in near real-time. gcloud compute ssh gives shell access, not logs. gcloud logging read queries past logs. gcloud app logs tail is for App Engine.

2
MCQeasy

An organization wants to connect their on-premises data center to Google Cloud with a dedicated 10 Gbps link. They require high availability and have budget for two physically diverse connections. Which solution should they choose?

A.Use Partner Interconnect with a single 10 Gbps connection.
B.Configure a single Dedicated Interconnect connection and use Cloud VPN as backup.
C.Provision two Cloud Dedicated Interconnect connections from diverse peering points.
D.Deploy a single HA VPN tunnel.
AnswerC

Two Dedicated Interconnect connections terminating at diverse peering points provide physically separate 10 Gbps paths, so a single facility or link failure does not drop connectivity. This satisfies both the dedicated bandwidth and high availability constraints.

Why this answer

Two Dedicated Interconnect connections from diverse peering points provide the required 10 Gbps dedicated bandwidth with high availability, since each connection is physically separate and can fail independently. This is the standard Google-recommended topology for production Dedicated Interconnect with redundancy.

Exam trap

PCA often tests whether candidates conflate 'high availability' with 'backup' and pick a single Dedicated Interconnect plus VPN, so the trap is missing that true HA for Dedicated Interconnect requires two physically diverse dedicated connections.

How to eliminate wrong answers

Option A is wrong because a single Partner Interconnect connection does not provide the redundancy required for high availability, and Partner Interconnect is delivered through a partner rather than a direct dedicated link. Option B is wrong because a single Dedicated Interconnect with Cloud VPN backup provides only one dedicated 10 Gbps path; the VPN backup is lower bandwidth and does not meet the two physically diverse dedicated connections requirement. Option D is wrong because a single HA VPN tunnel is an encrypted internet-based solution, not a dedicated 10 Gbps interconnect, and a single tunnel does not provide the required diversity.

3
Multi-Selectmedium

A media company runs a video transcoding service on GKE Standard. The service experiences sudden traffic spikes, and the operations team wants to ensure that the cluster can scale nodes automatically and that pods are rescheduled quickly when a node fails. The team also wants to monitor and alert on resource saturation. Which two actions should the cloud architect take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable cluster autoscaler on the node pools and set appropriate minimum and maximum node counts based on expected peak load.
B.Create a HorizontalPodAutoscaler based on CPU utilization for the transcoding deployment to add more pods when demand increases.
C.Enable Cloud CDN in front of the transcoding service to cache video segments and reduce load on the GKE pods.
D.Configure pod disruption budgets for the transcoding deployment to guarantee a minimum number of available pods during voluntary disruptions.
E.Set resource requests and limits on the transcoding pods so the scheduler and autoscaler can make accurate decisions about capacity and placement.
AnswersA, E

Cluster autoscaler adjusts the number of nodes in a node pool when pods cannot be scheduled due to insufficient resources, and it removes underutilized nodes when demand falls. Setting minimum and maximum counts bounds cost and capacity. This directly addresses automatic node scaling during traffic spikes and is a core reliability control for GKE workloads.

Why this answer

Automatic node scaling requires cluster autoscaler on the node pools, and it works correctly only when pods declare accurate resource requests and limits so the scheduler and autoscaler can size capacity. Together these ensure nodes are added during spikes and pods are placed and rescheduled efficiently. Pod disruption budgets, Cloud CDN, and HorizontalPodAutoscaler address different concerns and do not provide node-level scaling or rapid recovery from node failure.

Exam trap

The trap here is confusing pod-level scaling with node-level scaling, assuming that a HorizontalPodAutoscaler alone will add capacity when in fact cluster autoscaler is what provisions new nodes.

4
MCQmedium

You need to monitor the performance of a production Cloud Run service and set an alert when the p99 latency exceeds 500 ms over a 5-minute window. Which combination of Cloud Monitoring resources should you use?

A.Define an alerting policy using the metric 'run.googleapis.com/request_latencies' with a percentile aggregator and threshold condition
B.Create a log-based metric for latency and an alerting policy with a condition on the count of logs
C.Use Cloud Logging to export logs to BigQuery and run a scheduled query to check latency
D.Create an uptime check and set an alert on the check response time
AnswerA

Cloud Run exports request latency as a distribution metric, so a percentile aggregator is required to derive p99 rather than an average. Pairing that aggregator with a 500 ms threshold over a five-minute alignment window satisfies the stem's latency alerting condition.

Why this answer

Cloud Run exposes request latency as the built-in metric run.googleapis.com/request_latencies, which can be aggregated with a percentile aligner (e.g., 99th percentile) over a 5-minute window and used in a Cloud Monitoring alerting policy with a threshold of 500 ms. This is the native, lowest-latency path for p99 latency alerting on Cloud Run.

Exam trap

PCA often tests the misconception that log-based metrics or uptime checks can substitute for native latency metrics — the exam expects you to know that request_latencies with a percentile aggregator is the correct primitive for p99 alerting.

How to eliminate wrong answers

Option B is wrong because log-based metrics count log entries, not latency distributions — you cannot compute a p99 from log counts, and latency is already a first-class metric. Option C is wrong because exporting logs to BigQuery and running scheduled queries introduces significant delay and complexity, and logs do not contain structured latency percentiles suitable for real-time alerting. Option D is wrong because uptime checks measure availability and response time from external probes at a coarse interval, not the service's internal p99 request latency distribution.

5
MCQeasy

A small startup is deploying a new application on Google Cloud. They want to ensure that they can monitor the application's performance and receive alerts when certain thresholds are exceeded. They have limited operational staff and want a managed solution that requires minimal configuration. Which Google Cloud service should they use?

A.Cloud Logging with log sinks to BigQuery and custom scripts to analyze logs.
B.Cloud Monitoring with alerting policies based on metrics and log-based alerts.
C.Cloud Trace with custom instrumentation and manual analysis of trace data.
D.Cloud Profiler with continuous profiling and manual review of profiles.
AnswerB

Cloud Monitoring is a managed service that collects metrics, logs, and events from Google Cloud resources and applications. It allows you to create alerting policies that trigger notifications when metrics cross thresholds. It also supports log-based alerts for specific log patterns. With minimal configuration, the startup can set up dashboards and alerts without managing any infrastructure. This directly meets the requirement for a managed monitoring and alerting solution.

Why this answer

Cloud Monitoring is the managed service on Google Cloud for collecting metrics, creating dashboards, and setting up alerting policies. It integrates with many Google Cloud services and can send notifications via email, SMS, Slack, PagerDuty, and more. For a startup with limited staff, it requires minimal setup and no infrastructure management.

Log-based alerts extend its capabilities to specific log events. This makes it the ideal choice for monitoring performance and receiving alerts.

Exam trap

The trap here is confusing monitoring with logging, tracing, or profiling; only Cloud Monitoring provides alerting based on metrics and thresholds out of the box.

6
MCQmedium

Your company runs a stateful application on Compute Engine instances in a managed instance group (MIG). The application writes data to a persistent disk attached to each instance. You need to ensure that the application can automatically recover from a zone failure by recreating instances in another zone with their persistent disks. You also want to minimize data loss. Which configuration should you implement?

A.Create a zonal MIG in a single zone and use a snapshot schedule to back up persistent disks every hour to a multi-region bucket. In case of zone failure, manually restore the snapshots to new instances in another zone.
B.Create a regional MIG with instances distributed across multiple zones in the region, and configure each instance to use a regional persistent disk that is replicated across those zones.
C.Create a regional MIG and attach a standard persistent disk to each instance. Configure the MIG to recreate instances in other zones, and rely on the persistent disk's automatic replication across zones.
D.Create a zonal MIG and configure an autoscaler to add instances in other zones when the primary zone fails. Use a Cloud Storage bucket to store application data instead of persistent disks.
AnswerB

A regional MIG automatically distributes instances across zones and can recreate failed instances in other zones. Regional persistent disks are synchronously replicated across two zones in the same region, so if one zone fails, the disk can be attached to an instance in the other zone with minimal data loss. This combination provides automatic recovery and data redundancy.

Why this answer

A regional MIG spreads instances across multiple zones and can automatically recreate instances in healthy zones if one zone fails. Regional persistent disks are replicated across two zones, ensuring that data is available in another zone with minimal loss. Together, they provide automatic recovery and data redundancy.

Zonal MIGs cannot move instances across zones, and standard persistent disks are not replicated.

Exam trap

The trap here is assuming that a zonal MIG can automatically recreate instances in another zone or that standard persistent disks are replicated across zones.

7
MCQmedium

An organization is using Cloud Interconnect to connect their on-premises network to Google Cloud. They need to ensure 99.99% availability for their connection. Which configuration meets this requirement?

A.A single Partner Interconnect connection at 1Gbps
B.Two Dedicated Interconnect connections from different edge locations
C.A single Dedicated Interconnect connection at 10Gbps
D.High Availability VPN (HA VPN) with two gateways and four tunnels
AnswerB

Two Dedicated Interconnect connections terminating in different Google edge locations provide physically diverse paths, so a single edge or link failure cannot sever connectivity. This redundancy is what satisfies the 99.99% availability requirement, which a single connection or one edge location cannot guarantee.

Why this answer

To achieve 99.99% availability with Cloud Interconnect, Google Cloud requires two Dedicated Interconnect connections in different edge locations (or two Partner Interconnect connections in different metros). This redundant topology eliminates single points of failure at the interconnect level, meeting the SLA for 99.99% availability. A single connection, regardless of bandwidth, cannot meet this SLA.

Exam trap

PCA often tests SLA-to-topology mapping — candidates see '10Gbps' or 'HA VPN' and assume higher bandwidth or the word 'HA' guarantees 99.99%, missing that only redundant Interconnect connections in different edge locations meet the requirement.

How to eliminate wrong answers

Option A is wrong because a single Partner Interconnect connection at 1Gbps is a single point of failure and does not meet the 99.99% availability SLA — Partner Interconnect only reaches 99.99% with redundant connections in different metros. Option C is wrong because a single Dedicated Interconnect connection at 10Gbps, despite higher bandwidth, is still a single point of failure and only qualifies for 99.9% availability. Option D is wrong because HA VPN with two gateways and four tunnels provides 99.99% availability for VPN connectivity, but the question specifies Cloud Interconnect, and HA VPN is a different product that does not satisfy the Interconnect requirement.

8
MCQmedium

Your team runs a stateful analytics workload on a Managed Instance Group (MIG) of Compute Engine VMs. The VMs write intermediate results to local SSD scratch disks. During a recent incident, an autoscaling event terminated VMs and the intermediate data was lost, causing hours of recomputation. You need to change the deployment so that when a VM is terminated by the autoscaler, a shutdown script has enough time to flush the intermediate results to a Cloud Storage bucket before the VM is deleted. What should you do?

A.Replace the local SSD scratch disks with Persistent Disk volumes and rely on the autoscaler to detach them before deleting the VM.
B.Enable live migration on the MIG and set the autoscaler to scale in only during off-peak hours.
C.Set the MIG autoscaler's scale-in control to a longer cool-down period, and increase the instance template's minimum CPU utilization target.
D.Configure the instance template with a shutdown script and set the instance's shutdown duration metadata key to a value that gives the script enough time to flush data.
AnswerD

Compute Engine supports a per-instance shutdown duration, specified through the shutdown-duration metadata key, which extends the time the instance stays in the STOPPING state so a shutdown script can complete. Setting this on the instance template ensures every VM created by the MIG, including autoscaler-created ones, has enough time to flush intermediate results to Cloud Storage.

Why this answer

The shutdown duration metadata key is the supported mechanism to extend the STOPPING state so a shutdown script can finish. Because the MIG creates instances from the instance template, placing the key in the template guarantees that autoscaler-created VMs inherit the behavior. Autoscaler tuning and disk-type changes do not bound the time between the shutdown signal and instance deletion.

Exam trap

The trap here is assuming that autoscaler cool-down or scheduling controls how long a terminating VM stays alive, when only the shutdown duration setting actually extends that window.

9
MCQhard

During a load test, an application running on GKE experiences high latency and errors. You suspect the issue is due to insufficient cluster resources. Which gcloud command should you use to quickly check the current resource utilization of all nodes in the cluster?

A.gcloud container clusters describe
B.gcloud container clusters list
C.gcloud container clusters get-credentials
D.gcloud compute instances list
AnswerD

This command lists all compute instances, including GKE cluster nodes, and shows their status. While it does not display real-time CPU/memory usage, it quickly provides a list of nodes and their current state, which can help identify if nodes are down or overutilized based on status. However, for exact resource utilization, you would use kubectl top nodes.

Why this answer

None of the listed gcloud commands shows current resource utilization of GKE nodes. gcloud container clusters describe shows configuration, gcloud container clusters list lists clusters, gcloud container clusters get-credentials retrieves credentials, and gcloud compute instances list only lists instances and their status. The correct command for current CPU/memory utilization is kubectl top nodes.

Exam trap

Do not assume that gcloud container clusters describe or gcloud compute instances list shows resource utilization. They do not provide current CPU/memory metrics; use kubectl top nodes for that.

10
MCQhard

You are responsible for ensuring the reliability of a high-traffic web application running on Google Kubernetes Engine (GKE). You need to implement a monitoring strategy that alerts you when the application's error rate exceeds 1% over a 5-minute window. You want to minimize false positives and ensure alerts are actionable. What should you do?

A.Configure a log-based metric in Cloud Logging that counts log entries with severity ERROR, and create an alerting policy if the count exceeds a threshold.
B.Create a Cloud Monitoring alerting policy based on the HTTP load balancer's 5xx error rate metric, with a threshold of 1% and a duration of 5 minutes.
C.Use Prometheus to scrape application metrics and configure an alert in Prometheus Alertmanager for error rate > 1% over 5 minutes.
D.Create an uptime check in Cloud Monitoring that checks the application's health endpoint and alerts if it fails for 5 minutes.
AnswerB

The HTTP(S) load balancer exposes metrics such as request count and error count, which can be used to compute a 5xx error rate. An alerting policy with a 1% threshold over a 5-minute duration matches the requirement. Using the load balancer metric is reliable because it captures errors at the edge and is not affected by pod restarts or internal issues.

Why this answer

Using the HTTP(S) load balancer's 5xx error rate metric provides a direct measure of errors as seen by clients. An alerting policy with a 1% threshold over 5 minutes aligns with the requirement. This approach is managed, reliable, and minimizes false positives because it uses a consistent metric from the load balancer, not application logs or internal metrics.

Exam trap

The trap here is relying on log-based metrics or uptime checks, which may not accurately reflect the actual error rate experienced by users.

11
MCQeasy

To achieve a 99.999% availability SLA for a globally distributed application using Cloud Spanner, which configuration is required?

A.Multi-region instance configuration
B.Fine-grained access control
C.Single-region instance configuration
D.Customer-managed encryption keys (CMEK)
AnswerA

A multi-region instance configuration replicates data across regions with synchronous quorum writes, which is the only Spanner topology whose SLA reaches 99.999%. Regional configurations cap at 99.99%, so multi-region satisfies the availability constraint in the stem.

Why this answer

Cloud Spanner multi-region configuration provides 99.999% SLA. Single-region offers 99.99%. Fine-grained access control and customer-managed encryption keys (CMEK) do not affect availability SLA.

12
Multi-Selectmedium

An organization needs to implement a change management process for a mission-critical application on GKE. They want to validate performance before full rollout and be able to roll back quickly. Which THREE practices should they adopt? (Choose THREE.)

Select 3 answers
A.Deploy changes directly to production
B.Implement canary deployments with traffic splitting
C.Use feature flags to enable/disable features dynamically
D.Manually monitor and roll back if issues appear
E.Define automated rollback policies in Cloud Deploy
AnswersB, C, E

Canary deployments with traffic splitting route a small percentage of live requests to the new revision, letting the team measure real performance before full rollout. If metrics degrade, shifting traffic back to the stable revision provides near-instant rollback.

Why this answer

Option B is correct because canary deployments with traffic splitting (e.g., via GKE Ingress, Anthos Service Mesh, or Istio VirtualService weights) let the team expose a new version to a small percentage of traffic, validating performance on real workloads before full rollout. Option C is correct because feature flags decouple deployment from release, allowing features to be toggled dynamically without redeploying, which supports fast rollback by disabling a flag rather than reverting a build. Option E is correct because Cloud Deploy supports automated rollback policies tied to rollout failures or verification results, enabling rapid, deterministic rollback of a bad release.

Option A is wrong because deploying directly to production bypasses validation and contradicts the requirement to test performance before full rollout. Option D is wrong because manual monitoring and rollback is slow, error-prone, and does not meet the goal of rolling back quickly compared to automated policies.

13
MCQmedium

A company is planning a phased migration of their on-premises database to Cloud SQL. They want to minimize downtime and ensure data consistency. Which approach should they use?

A.Use Database Migration Service (DMS)
B.Lift and shift the database server to Compute Engine
C.Export the database to a SQL dump file and import into Cloud SQL
D.Use VM migration to move the database server
AnswerA

Database Migration Service performs continuous replication from the source into Cloud SQL, keeping the target synchronised until cutover. This minimises downtime and preserves consistency during the phased migration, unlike a one-off dump and load, which requires pausing writes.

Why this answer

Database Migration Service (DMS) supports continuous replication with minimal downtime. Export and import involves downtime. Lift-and-shift is not a GCP service.

VM migration is for servers, not databases.

14
MCQhard

An application running on Compute Engine is experiencing increased latency. You suspect a network bottleneck due to high egress traffic. Which gcloud command can you use to quickly check the network egress traffic for a specific VM instance?

A.gcloud logging read 'resource.type=gce_instance AND jsonPayload.egress_bytes'
B.gcloud compute instances list --format='value(networkInterfaces[0].networkIP)'
C.gcloud compute instances get-serial-port-output
D.gcloud monitoring metrics list
AnswerA

gcloud logging read with the specified filter queries Cloud Logging for egress bytes logs from Compute Engine instances, making it the correct choice.

Why this answer

The gcloud logging read command allows you to query Cloud Logging for specific log entries. For a Compute Engine instance, the resource type is gce_instance. You can filter for egress bytes by using the query 'resource.type=gce_instance AND jsonPayload.egress_bytes'. This will return log entries containing egress bytes information, assuming your VM is configured to send these logs (e.g., via the monitoring agent or VPC flow logs). This is the quickest way among the given options to check network egress traffic for a specific VM using a native gcloud command.

Option B is incorrect: gcloud compute instances list only displays the network IP of instances, not egress traffic metrics.

Option C is incorrect: gcloud compute instances get-serial-port-output shows the serial console output, which does not include network traffic data.

Option D is incorrect: gcloud monitoring metrics list just lists available metrics but does not retrieve the actual traffic data for a specific instance. To get the data you would need to use gcloud monitoring metric descriptors or gcloud monitoring dashboards, but the command as given does not return traffic.

Thus, A is the best choice.

Exam trap

Students may think that Cloud Monitoring is the only way to view metrics, but Cloud Logging can also be used to query specific data like egress bytes if logs are collected. They might also incorrectly choose D because monitoring sounds relevant, but the command 'gcloud monitoring metrics list' only lists metric descriptors, not actual data.

15
MCQhard

Your organization operates a multi-project Google Cloud environment. A security team requires that any new Compute Engine instance created in the production folder must have OS Login enabled and must not use project-wide SSH keys. You want to enforce this centrally with the least operational overhead and ensure that non-compliant creation attempts are denied. What should you do?

A.Create an Organization Policy constraint for compute.requireOsLogin and compute.disableProjectSshKeys, and apply them at the production folder.
B.Create an Organization Policy constraint for compute.requireOsLogin and compute.skipDefaultNetworkCreation, and apply it at the production folder.
C.Apply a custom IAM deny policy on compute.instances.create for all principals except the security team, and require them to create instances on behalf of others.
D.Grant the security team the Compute Security Admin role and schedule a Cloud Scheduler job that audits instances with project-wide SSH keys.
AnswerA

Organization Policy constraints compute.requireOsLogin and compute.disableProjectSshKeys are designed exactly for this requirement. Applying them at the production folder enforces OS Login and blocks project-wide SSH keys for all projects under that folder, denying non-compliant instance creation without per-project scripting or IAM churn.

Why this answer

Organization Policy constraints applied at the folder level enforce configuration rules across all descendant projects. compute.requireOsLogin forces OS Login for instances, and compute.disableProjectSshKeys blocks adding project-wide SSH keys. Applying both at the production folder denies non-compliant creation centrally with minimal operational overhead, unlike audits or IAM restrictions that do not enforce the desired configuration.

Exam trap

The trap here is confusing an audit-and-alert approach or an IAM restriction with preventive configuration enforcement, which only Organization Policy constraints provide.

16
MCQmedium

Your team is deploying a new version of a microservices application on Google Kubernetes Engine (GKE). You want to gradually shift traffic to the new version while monitoring key performance indicators (KPIs) such as error rate and latency. If KPIs degrade, you need to automatically roll back. Which approach should you use?

A.Use GKE rolling updates with maxSurge and maxUnavailable set to 25%, and monitor KPIs manually.
B.Implement a canary deployment using Istio with Flagger, which automates traffic shifting and rollback based on Prometheus metrics.
C.Configure a GKE Ingress with two backends and use traffic splitting based on weights, then manually adjust weights based on monitoring.
D.Use Blue/Green deployment by creating a second deployment and switching the Service selector to the new version after manual testing.
AnswerB

Flagger is a progressive delivery tool that integrates with Istio to automate canary releases. It gradually shifts traffic, monitors Prometheus metrics for KPIs, and automatically rolls back if metrics breach thresholds. This matches the requirement for automated rollback based on KPIs.

Why this answer

Flagger with Istio automates canary deployments by incrementally shifting traffic, analyzing Prometheus metrics, and rolling back automatically if KPIs degrade. This provides safe, gradual rollout with minimal manual intervention. Other options either lack automation, do not support gradual traffic shifting, or require manual monitoring.

Exam trap

The trap here is confusing rolling updates with canary deployments; rolling updates replace pods but do not control traffic splitting or provide automated rollback based on metrics.

17
Multi-Selectmedium

You are responsible for operations reliability of a production service running on Google Cloud. The service is deployed on GKE and exposes an external HTTPS endpoint through an external Application Load Balancer. You need to implement monitoring that detects when the service is unhealthy from the user's perspective and alerts the on-call team. (Choose two.)

Select 2 answers
A.Create an alerting policy on the Application Load Balancer's request count metric to fire when traffic drops below a static threshold.
B.Enable Cloud Trace on the GKE workloads and alert when the number of spans per minute exceeds a fixed value.
C.Set up an alerting policy on the load balancer's 5xx error rate and on backend latency, with thresholds tied to the service level objective.
D.Configure a log-based alert on GKE node system logs for the keyword 'OOMKilled'.
E.Create an uptime check in Cloud Monitoring that targets the external HTTPS URL and verifies the expected response code and content.
AnswersC, E

Load balancer 5xx rates and backend latency metrics capture server-side failures and performance degradation as seen at the edge. Alerting on these signals against SLO-derived thresholds detects when users experience errors or slow responses. Combined with an uptime check, this provides both external probing and internal telemetry for reliable detection.

Why this answer

Detecting user-facing unhealthiness requires probing the service as users reach it and monitoring the error and latency signals that reflect their experience. An uptime check from multiple locations validates the public endpoint and response content, while alerting on load balancer 5xx rates and backend latency ties detection to SLO thresholds. Together they catch both total outages and degraded performance.

Exam trap

The trap here is choosing internal resource metrics, such as OOMKilled logs or span counts, as the primary health signal instead of external probes and edge-level error and latency metrics.

18
MCQeasy

A startup runs a customer-facing web application on Cloud Run. The operations team needs to know when the service's request latency exceeds a threshold so they can respond before users complain. They want to be notified by email and also want a record of the incident for later review. Which Google Cloud service should they use to define the alerting policy?

A.Cloud Trace analysis reports scheduled to run daily and emailed to the operations team.
B.Cloud Monitoring alerting policies with a notification channel for email.
C.Cloud Run revision traffic splitting combined with a health check endpoint that returns an error when latency is high.
D.Cloud Logging log-based alerts with a notification channel for email.
AnswerB

Cloud Monitoring alerting policies evaluate metrics such as request latency against thresholds and trigger notifications through configured channels like email. They also record incidents and their state transitions, giving the team both immediate notification and a historical record. This is the native, integrated way to alert on Cloud Run latency metrics.

Why this answer

Cloud Monitoring alerting policies are designed to evaluate metrics like request latency against thresholds and to notify through channels such as email. They also track incidents over time, satisfying the need for both immediate notification and a reviewable record. Log-based alerts, Cloud Trace reports, and traffic splitting serve different purposes and do not provide threshold-based latency alerting with email notification.

Exam trap

The trap here is choosing log-based alerting for a numeric metric threshold, when log-based alerts fire on matching log entries rather than on values crossing a latency threshold.

19
MCQmedium

Your company runs a multi-region Cloud Spanner instance for a global financial application. The SLA requirement is 99.999% availability. You need to ensure that the database remains available during a regional outage. What configuration should you use?

A.Use a dual-region configuration with two regions but only one for writes.
B.Use a single-region configuration with a read replica in another region.
C.Use a multi-region configuration (e.g., nam3) with automatic replication across multiple regions.
D.Configure a single-region instance and create periodic backups to restore in another region.
AnswerC

Multi-region configurations such as nam3 replicate synchronously across regions, so a single regional outage does not interrupt reads or writes. This satisfies the 99.999% SLA, which a regional or single-region instance cannot deliver during regional failure.

Why this answer

Cloud Spanner multi-region configurations (e.g., nam3, eur3) automatically replicate data across regions within a continent. They provide 99.999% availability SLA. A single-region configuration offers 99.99% SLA.

Read replicas (as in Cloud SQL) are not a concept in Spanner. Multi-region configs use multiple read-write regions.

20
MCQeasy

Your company has a service running on Google Kubernetes Engine (GKE) that experiences occasional spikes in traffic. You need to ensure that the service remains available during these spikes by automatically scaling the number of pods based on CPU utilization. You also want to minimize cost by scaling down when traffic decreases. Which Kubernetes resource should you configure?

A.A Cluster Autoscaler with a node pool that has a minimum of 2 nodes and a maximum of 10 nodes.
B.A VerticalPodAutoscaler with a target CPU utilization of 80%.
C.A HorizontalPodAutoscaler with a target CPU utilization of 80% and a minimum of 2 replicas and a maximum of 10 replicas.
D.A PodDisruptionBudget with minAvailable set to 80%.
AnswerC

HorizontalPodAutoscaler automatically scales the number of pods in a Deployment or ReplicaSet based on observed CPU utilization or other metrics. Setting a target CPU utilization of 80% ensures that when average CPU exceeds 80%, more pods are added, and when it drops, pods are removed down to the minimum. This directly addresses traffic spikes and cost optimization.

Why this answer

HorizontalPodAutoscaler is the Kubernetes resource designed to scale the number of pods based on metrics like CPU utilization. By setting a target CPU utilization and replica bounds, it can automatically add pods during spikes and remove them when demand drops, ensuring availability and cost efficiency. The other options either adjust resources vertically, scale nodes, or manage disruptions, none of which directly scale pods based on CPU.

Exam trap

The trap here is confusing VerticalPodAutoscaler with HorizontalPodAutoscaler; vertical scaling adjusts resources per pod and does not add pods.

21
MCQmedium

Your team is responsible for a production service running on Google Cloud. You need to define Service Level Objectives (SLOs) and monitor them using Cloud Monitoring. You want to be alerted when the service's error budget is being consumed too quickly. Which approach should you take?

A.Use Cloud Monitoring's SLO monitoring to define an SLO with a 99.9% availability target, and create an alerting policy based on the burn rate of the error budget.
B.Set up a log-based metric that counts errors, and create an alert when the count exceeds a certain number within a 5-minute window.
C.Configure a dashboard in Cloud Monitoring that displays the error rate and set up a cron job to check the dashboard every hour and send an email if the error rate is high.
D.Create an alerting policy that triggers when the error rate exceeds a threshold of 1% over a 1-hour window.
AnswerA

Cloud Monitoring allows you to define SLOs and then create alerting policies that trigger based on the burn rate of the error budget. This is the recommended practice for alerting on SLO violations. You can set thresholds for burn rates over different windows (e.g., 2% in 1 hour) to detect both fast and slow burns, enabling proactive response before the budget is exhausted.

Why this answer

Cloud Monitoring's SLO monitoring feature allows you to define SLOs and then create alerting policies based on error budget burn rates. This is the most effective way to alert on SLO violations because it considers both the error rate and the time window, and it can detect when the budget is being consumed too quickly. Other methods lack the direct integration with SLOs and error budgets.

Exam trap

The trap here is assuming that a simple error rate threshold is sufficient, without considering the error budget burn rate and the SLO target.

22
MCQmedium

A company is designing a disaster recovery (DR) plan for their Cloud SQL for PostgreSQL instance. They need to recover the database to a specific point in time within the last 7 days, with a Recovery Point Objective (RPO) of less than 1 hour. Which feature should they use?

A.Exporting the database daily to Cloud Storage
B.Point-in-time recovery (PITR)
C.Failover replica
D.Automated backups only
AnswerB

PITR continuously archives transaction logs, letting Cloud SQL for PostgreSQL restore to any second within the retention window. This meets the sub-one-hour RPO and the seven-day recovery target, unlike daily automated backups, which cap recovery granularity at 24 hours.

Why this answer

Point-in-time recovery (PITR) for Cloud SQL for PostgreSQL lets you restore an instance to any specific timestamp within a configurable retention window (up to 7 days), using write-ahead log (WAL) archiving combined with automated backups. This directly satisfies the requirement to recover to a specific point in time within the last 7 days with an RPO under 1 hour, because WAL segments are continuously shipped and enable granular recovery. Automated backups alone only allow restore to the backup's snapshot time, which would not meet a sub-hour RPO.

Exam trap

PCA often tests the difference between HA (failover replica) and DR (PITR/backups) — the trap is choosing 'failover replica' because it sounds resilient, when it actually propagates logical corruption instead of enabling recovery to an earlier point.

How to eliminate wrong answers

Option A is wrong because daily exports to Cloud Storage only capture a point-in-time snapshot once per day, yielding an RPO of up to 24 hours and no ability to recover to an arbitrary point within the day. Option C is wrong because a failover replica provides high availability (automatic failover to a standby in another zone) but does not enable point-in-time recovery to an earlier timestamp — it mirrors the current state, including any logical corruption. Option D is wrong because automated backups alone restore only to the time the backup was taken (typically daily), which cannot meet a sub-1-hour RPO or a specific point-in-time requirement.

23
MCQmedium

Your company runs a production microservices application on GKE Standard. The operations team wants to be notified when any pod in the cluster is repeatedly restarting, indicating a potential CrashLoopBackOff. They want to use Cloud Monitoring to create an alert that fires when a container restarts more than 5 times in a 10-minute window. Which metric should they use as the basis for the alerting policy?

A.kubernetes.io/container/cpu/core_usage_time
B.kubernetes.io/container/restart_count
C.kubernetes.io/pod/status
D.logging.googleapis.com/user/restart_count
AnswerB

This metric is a cumulative counter that tracks the number of times a container has restarted. By using a rate or delta alignment over a 10-minute window, you can detect when restarts exceed a threshold of 5, directly matching the requirement. It is the standard metric for container restarts in GKE and is available in Cloud Monitoring without additional setup.

Why this answer

The kubernetes.io/container/restart_count metric directly tracks container restarts and is the correct choice for alerting on repeated restarts. It is a cumulative counter, so you must apply a rate or delta alignment to detect increases over time. Other metrics like CPU usage or pod status do not provide restart counts, and log-based metrics require extra setup and may be less reliable.

Exam trap

The trap here is assuming that pod status or CPU metrics can indicate restarts, when only the dedicated restart count metric provides the necessary data.

24
MCQmedium

A retail company runs a web application on Compute Engine instances behind an external HTTP(S) load balancer. During a flash sale, the operations team notices that the load balancer is returning HTTP 502 errors for a subset of requests. The backend service health checks are passing, and the instances are not under heavy CPU load. The team wants to identify the root cause quickly and prevent recurrence. Which action should they take first?

A.Configure a Cloud Armor security policy to block traffic from suspicious IP addresses.
B.Review the load balancer's logs in Cloud Logging for HTTP 502 status codes and examine the backend service's health check logs.
C.Enable Cloud CDN on the backend service to cache static content and reduce load on the instances.
D.Increase the backend service's timeout setting from 30 seconds to 60 seconds.
AnswerB

The first step in troubleshooting is to gather data. Cloud Logging captures load balancer logs that include the status code, backend instance, and error details. By filtering for 502 errors, the team can identify patterns such as a specific instance or a particular request path. Examining health check logs can reveal if health checks are flapping or if instances are being marked unhealthy intermittently. This data-driven approach identifies the root cause before making changes.

Why this answer

The most effective first step is to use Cloud Logging to examine load balancer logs and health check logs. These logs provide detailed information about the 502 errors, including which backend instances are involved and any error messages. By analyzing this data, the team can pinpoint the root cause, such as a misconfigured backend, an application error, or a network issue.

Only after identifying the cause should they take corrective action to prevent recurrence.

Exam trap

The trap here is jumping to configuration changes like increasing timeouts or enabling CDN without first diagnosing the actual cause of the 502 errors.

25
Multi-Selectmedium

Your company is deploying a new application on Google Cloud and needs to ensure that it can meet a 99.9% availability SLA. You are designing the architecture for high availability. Which two practices should you implement? (Choose two.)

Select 2 answers
A.Implement health checks and autohealing for managed instance groups.
B.Use a single global load balancer to distribute traffic across all instances.
C.Deploy the application across multiple zones within a single region.
D.Use a single zone with a high-capacity machine type to reduce complexity.
E.Store application state on a local SSD attached to each instance.
AnswersA, C

Health checks and autohealing ensure that unhealthy instances are automatically recreated. This maintains the desired capacity and availability of the application. When combined with multi-zone deployment, autohealing helps recover from instance failures quickly, contributing to meeting high availability SLAs.

Why this answer

To achieve high availability, you should deploy across multiple zones within a region to survive zone failures and implement health checks with autohealing to automatically recover from instance failures. These two practices together ensure that the application remains available even when individual instances or zones fail, helping to meet a 99.9% SLA.

Exam trap

The trap here is assuming that a global load balancer alone provides high availability, when it must be paired with multi-zone deployment and autohealing to be effective.

26
MCQmedium

Your team manages a production web application on Compute Engine behind an external Application Load Balancer. During a recent incident, the load balancer's backend service marked all instances as unhealthy because the health check endpoint returned HTTP 200 but the application was actually in a degraded state. You need Cloud Monitoring to alert the operations team when the application's error rate exceeds 5% over a 5-minute window. You also need to ensure that the alert does not fire during planned maintenance windows. Which approach should you take?

A.Use Cloud Trace to sample requests and create an alerting policy based on the latency of traces that return errors, triggering when error traces exceed 5% of total traces.
B.Create an uptime check that sends HTTP requests to the application's health endpoint every minute and alerts when the check fails from more than one region.
C.Configure a custom health check on the load balancer that returns HTTP 500 when the application is degraded, and create an alerting policy on the backend service's unhealthy instance count.
D.Create a log-based metric that counts HTTP 5xx responses from the load balancer logs, then create an alerting policy on that metric with a threshold of 5% error rate and configure a maintenance window for planned downtime.
AnswerD

Log-based metrics derive values from log entries, and the load balancer logs include status details for each request. By counting 5xx responses and dividing by total requests, you can compute an error rate. Alerting policies can use such metrics with threshold conditions, and maintenance windows suppress alerts during planned downtime. This directly addresses the requirement to monitor application-level errors rather than infrastructure health, and respects maintenance periods.

Why this answer

The requirement is to alert on application error rate exceeding 5% over 5 minutes, with suppression during maintenance. A log-based metric from load balancer logs captures every request's status, enabling an accurate error rate calculation. Alerting policies on such metrics support threshold conditions and maintenance windows.

The other options either alter health checks in a harmful way, rely on uptime checks that miss application-level errors, or use sampled trace data that is not representative.

Exam trap

The trap here is assuming that a health check endpoint returning HTTP 200 means the application is healthy and that uptime checks or health check status can measure error rate.

27
MCQhard

Your company runs a production application on Compute Engine instances behind a managed instance group (MIG). You need to perform a rolling update with canary testing, gradually shifting traffic to the new version only if performance metrics are healthy. Which approach should you use?

A.Use Cloud Deploy with a deployment strategy that includes a canary phase and automated verification
B.Create a new MIG with the new template and use a Cloud Load Balancer's traffic splitting
C.Manually update each instance by SSH'ing and running a script
D.Use gcloud compute instance-groups managed rolling-action start-update with a maxSurge of 0
AnswerA

Cloud Deploy orchestrates the progressive rollout across the MIG, defining a canary phase that shifts a percentage of traffic, then runs automated verification against your defined metrics before promoting or rolling back. This satisfies the requirement to advance only when performance stays healthy.

Why this answer

Cloud Deploy is Google Cloud's managed continuous delivery service that natively supports progressive rollout strategies including canary deployments with automated verification against SLOs or custom metrics. It integrates with Cloud Monitoring to pause or roll back a rollout if health checks fail, which directly satisfies the requirement of gradually shifting traffic only when performance metrics are healthy. This is the intended Google-recommended pattern for canary releases on GCE MIGs.

Exam trap

The trap here is assuming that a MIG rolling update alone provides canary behavior; in reality, rolling updates replace instances without traffic-based verification, so candidates who pick the gcloud rolling-action option miss the automated canary requirement.

How to eliminate wrong answers

Option B is wrong because a second MIG plus load balancer traffic splitting requires manual orchestration of the split percentages and does not provide automated verification or rollback based on metrics. Option C is wrong because SSH-based manual updates are not rolling, not canary, and provide no automated health gating. Option D is wrong because 'rolling-action start-update' performs an in-place rolling replacement of instances but does not perform canary traffic shifting or metric-based verification — maxSurge of 0 also forces downtime-style replacement.

28
Multi-Selecthard

A company is designing a highly available architecture for a web application using Google Cloud. They need to ensure that the application remains available even if an entire Google Cloud region experiences an outage. Which THREE components should they include in their architecture? (Choose THREE.)

Select 3 answers
A.Cloud Spanner multi-region configuration
B.Cloud SQL with a cross-region read replica
C.Global external HTTP(S) load balancer
D.Cloud CDN
E.Regional managed instance groups in multiple regions
AnswersA, C, E

Cloud Spanner multi-region configuration synchronously replicates data across regions with strong consistency, so the database survives a full regional outage. This satisfies the stem's requirement that the application remain available despite an entire Google Cloud region failing.

Why this answer

Option A (Cloud Spanner multi-region configuration) is correct because a multi-region instance replicates data synchronously across regions with 99.999% availability, so the database survives a full regional outage without data loss. Option C (Global external HTTP(S) load balancer) is correct because it is a global anycast service that routes users to the nearest healthy backend and automatically fails over to backends in other regions when one region becomes unavailable. Option E (Regional managed instance groups in multiple regions) is correct because deploying regional MIGs in at least two regions provides compute capacity that keeps serving traffic when one region fails, and it pairs with the global load balancer for automatic failover.

Option B is not sufficient because Cloud SQL cross-region read replicas are asynchronous and require manual or scripted promotion, so they do not provide automatic, zero-data-loss regional failover. Option D is not correct because Cloud CDN only caches content at edge locations; it improves latency and offloads origin traffic but does not by itself provide regional failover for dynamic application workloads.

29
MCQeasy

A company wants to connect their on-premises network to Google Cloud with a 99.99% SLA using encrypted tunnels over the public internet. Which connectivity solution should they choose?

A.HA VPN
B.Standard VPN with single tunnel
C.Partner Interconnect
D.Dedicated Interconnect
AnswerA

HA VPN provides two tunnels across two interfaces, delivering a 99.99% availability SLA when configured with two external IP addresses. It encrypts traffic over the public internet, matching both the SLA and encryption constraints in the stem.

Why this answer

HA VPN provides a 99.99% SLA when configured with two VPN gateways and four tunnels over the public internet. Dedicated Interconnect is a private connection with higher bandwidth but not over the public internet. Partner Interconnect uses a partner's network, not the public internet.

Standard VPN does not offer a 99.99% SLA.

30
MCQeasy

A startup deploys a containerized web application on Cloud Run. They want to release a new revision to a small percentage of users before promoting it to all traffic, and they need the ability to roll back instantly if errors increase. Which Cloud Run feature should they use?

A.Use Cloud Deploy with a canary deployment strategy targeting the Cloud Run service and rely on its automatic rollback on failure.
B.Enable session affinity on the Cloud Run service and deploy the new revision so existing sessions stay on the old revision.
C.Create a second Cloud Run service for the new version and use a global external HTTP(S) load balancer with weighted backends to split traffic.
D.Deploy the new revision with --no-traffic, then use traffic splitting to send a percentage of requests to it, and adjust or roll back by changing the traffic allocation.
AnswerD

Cloud Run supports deploying a revision without traffic and then splitting traffic by percentage across revisions. This enables a canary release to a small share of users, and rollback is immediate by routing all traffic back to the previous revision. It requires no extra infrastructure and uses built-in revision management.

Why this answer

Cloud Run revisions are immutable and traffic can be split by percentage across them. Deploying with no traffic, then assigning a small percentage, implements a canary, and setting the previous revision to 100 percent rolls back instantly. This built-in capability avoids external load balancers or delivery pipelines for a straightforward canary and rollback.

Exam trap

The trap here is assuming a load balancer or a separate service is required for canary releases, when Cloud Run traffic splitting already provides percentage-based routing and instant rollback.

31
MCQeasy

Your company has a Service Level Objective (SLO) of 99.9% availability for a web application running on Google Cloud. You want to create an alert that notifies the on-call team when the error budget is being consumed too quickly. Which Google Cloud service should you use?

A.Cloud Logging with a log-based metric and an alerting policy.
B.Cloud Monitoring with an SLO and a burn rate alert.
C.Cloud Monitoring with an alerting policy based on a metric threshold for error rate.
D.Cloud Trace with latency thresholds and alerts.
AnswerB

Cloud Monitoring allows you to define an SLO and create burn rate alerts. Burn rate alerts notify when the error budget is being consumed at a rate that would exhaust it faster than desired. This directly addresses the requirement to alert on rapid error budget consumption.

Why this answer

Cloud Monitoring supports defining SLOs and creating burn rate alerts. Burn rate alerts fire when the rate of error budget consumption is too high, helping teams respond before the budget is exhausted. This is the native, recommended way to alert on SLO violations.

Exam trap

The trap here is thinking that a simple error rate threshold alert is sufficient, but it does not account for the SLO and burn rate, which are essential for proactive alerting.

32
Multi-Selectmedium

A company wants to set up monitoring and alerting for their application running on GKE. They need to receive alerts via email and also trigger an automated remediation workflow. Which TWO components should they use? (Choose two.)

Select 2 answers
A.Notification channels (email)
B.Alerting policies
C.Cloud Shell
D.Cloud Logging
E.Pub/Sub
AnswersA, B

Notification channels define the delivery mechanism, so an email channel satisfies the requirement to receive alerts by email. Alerting policies detect the condition, but the channel is what actually routes the notification to the recipient's inbox.

Why this answer

Option A, Notification channels (email), is correct because in Cloud Monitoring a notification channel defines the destination and delivery mechanism for alerts, and an email-type channel is exactly what is needed to receive alert notifications via email. Option B, Alerting policies, is correct because alerting policies define the conditions (metrics, thresholds, filters) that determine when an alert fires and which notification channels are used, making them the core component for monitoring and alerting on the GKE application. Together, an alerting policy detects the condition and routes it to the email notification channel, satisfying the email alerting requirement.

Option C, Cloud Shell, is just an interactive command-line environment and provides no monitoring or alerting capability. Option D, Cloud Logging, stores and queries logs but does not itself define alert conditions or deliver email notifications. Option E, Pub/Sub, can be used as a notification channel for automation, but it is not required to receive email alerts and is not one of the two components needed for the stated email alerting requirement.

Exam trap

PCA often tests the distinction between the alerting policy (the 'what/when') and the notification channel (the 'where'), and candidates mistakenly pick Pub/Sub or Cloud Logging as the alerting component instead of recognizing them as transport/storage layers.

33
MCQhard

A financial services company runs a payment API on Compute Engine behind an internal passthrough Network Load Balancer. The compliance team requires that all administrative actions on the project be attributable to a named human, that production changes be reviewed before taking effect, and that no single engineer can delete the production database. Which combination of Google Cloud controls should the cloud architect implement?

A.Require all engineers to use hardware security keys for two-factor authentication, enable Identity-Aware Proxy for SSH access to instances, and create an alerting policy that notifies the security team when the database is deleted.
B.Assign least-privilege predefined roles to engineers, require all production changes to go through a CI/CD pipeline that uses a dedicated service account in a separate project, and protect the production database with a resource-level deny policy and separation of duties.
C.Grant the Project Owner role to all senior engineers, enable Cloud Audit Logs for Admin Activity, and require them to use a shared break-glass account for emergency changes.
D.Enable VPC Service Controls around the production project, grant engineers the Editor role, and configure Cloud Logging sinks to export audit logs to a separate project for long-term retention.
AnswerB

Least-privilege roles limit what each engineer can do, and a pipeline with a dedicated service account in a separate project ensures changes are reviewed and executed by automation rather than directly by humans. A deny policy on the database prevents even privileged users from deleting it, and separation of duties keeps one person from both proposing and approving. This gives attribution, review, and protection.

Why this answer

The compliance demands map to preventive controls: least privilege to limit permissions, a reviewed CI/CD pipeline with a dedicated service account in a separate project to enforce change review and attribution, and a resource-level deny policy plus separation of duties to stop any one engineer from deleting the production database. Detective measures like audit logging and alerting are useful but insufficient on their own. The combination of IAM restrictions, automation, and deny policies satisfies all three requirements.

Exam trap

The trap here is treating detective controls such as audit log exports or deletion alerts as sufficient, when the compliance requirements demand preventive controls that stop unauthorized or unreviewed actions before they occur.

34
Multi-Selectmedium

You are designing a disaster recovery plan for a critical application running on GKE. You need to back up the cluster's state and application data. Which TWO services should you use together? (Choose 2)

Select 2 answers
A.Velero (formerly Heptio Ark)
B.Pub/Sub
C.Cloud SQL
D.Filestore
E.Cloud Storage
AnswersA, E

Velero backs up Kubernetes cluster state and persistent volume data, capturing the GKE resources and application data the plan requires. It satisfies the cluster-state constraint by exporting API objects and volume snapshots together for later restore.

Why this answer

Velero (A) is correct because it is the standard open-source tool for backing up and restoring Kubernetes cluster state, including namespaces, deployments, services, and persistent volume snapshots, making it ideal for GKE disaster recovery. Cloud Storage (E) is correct because Velero stores its backups and volume snapshots in an object storage backend, and Google Cloud Storage is the native, durable, and highly available GCS bucket option for GKE environments. Together, Velero orchestrates the backup and restore operations while Cloud Storage provides the persistent, off-cluster repository for those backups.

Pub/Sub (B) is a messaging service, not a backup or storage solution, so it does not back up cluster state or application data. Cloud SQL (C) is a managed relational database service that could host application data but does not back up GKE cluster state, and Filestore (D) is a managed NFS file share for persistent volumes, not a backup repository for cluster-wide state.

Exam trap

The trap is picking GCP-native services like Cloud SQL or Filestore because they sound data-related, when the question is specifically about backing up GKE cluster state and application data, which requires a Kubernetes-aware tool plus object storage.

35
MCQmedium

Your company plans to connect an on-premises data center to Google Cloud with a Dedicated Interconnect. You need to ensure high availability for the connection. What is the minimum configuration required to meet a 99.99% SLA for Dedicated Interconnect?

A.Two Dedicated Interconnect circuits, each in a different edge availability domain, with a Cloud Router for each connection
B.A single Dedicated Interconnect circuit with a Cloud Router configured for BGP advertisements
C.One Dedicated Interconnect circuit and one Partner Interconnect connection as a backup
D.One Dedicated Interconnect circuit with two VLAN attachments on the same circuit
AnswerA

Two circuits in separate edge availability domains give physical path redundancy, and a Cloud Router per connection satisfies the topology requirement for the 99.99% Dedicated Interconnect SLA. A single circuit or shared router only reaches 99.9%.

Why this answer

To achieve the 99.99% availability SLA for Dedicated Interconnect, Google requires at least two Dedicated Interconnect connections in two different edge availability domains (metro availability zones), each terminating on a separate Cloud Router. This topology ensures that a single circuit or router failure does not take down connectivity. The 99.99% SLA is explicitly tied to this redundant, diverse-path configuration.

Exam trap

The trap is assuming that any two connections (e.g., one Dedicated plus one Partner, or two VLAN attachments on one circuit) satisfy the redundancy requirement; the exam expects you to know that the 99.99% SLA specifically requires two Dedicated Interconnect circuits in different edge availability domains with separate Cloud Routers.

How to eliminate wrong answers

Option B is wrong because a single Dedicated Interconnect circuit with one Cloud Router provides only the 99.9% SLA — there is a single point of failure at both the circuit and router level. Option C is wrong because mixing one Dedicated Interconnect with one Partner Interconnect does not meet the documented topology for the 99.99% Dedicated Interconnect SLA; the SLA tiers are defined per product and per redundancy configuration, and this hybrid does not qualify. Option D is wrong because two VLAN attachments on the same physical circuit still share the same circuit and edge availability domain, so a circuit failure takes down both attachments — this is not redundancy at the physical layer.

36
MCQmedium

A company wants to connect their on-premises data center to Google Cloud with a dedicated private connection that provides 99.99% availability and supports up to 100 Gbps bandwidth. They have a colocation facility near a Google Cloud region. Which connectivity option should they choose?

A.Partner Interconnect
B.Direct Peering
C.Dedicated Interconnect
D.HA VPN
AnswerC

Dedicated Interconnect provides a direct physical link between the on-premises network and Google's edge at a colocation facility, delivering the 99.99% availability and up to 100 Gbps capacity the stem requires. Partner Interconnect and Cloud VPN cannot meet those combined bandwidth and SLA constraints.

Why this answer

Dedicated Interconnect provides a direct physical connection between the customer's colocation facility and Google's network, supporting up to 100 Gbps per link (with 10 Gbps or 100 Gbps circuits) and offering a 99.99% SLA when configured with redundant connections. Because the company already has a colocation facility near a Google region, Dedicated Interconnect is the correct fit for the stated bandwidth and availability requirements.

Exam trap

The trap is confusing Dedicated Interconnect with Partner Interconnect; candidates may pick Partner because it also uses a colocation facility, but only Dedicated Interconnect meets the 100 Gbps and direct-connection requirements.

How to eliminate wrong answers

Option A is wrong because Partner Interconnect goes through a supported service provider and typically supports lower bandwidth tiers (up to 50 Gbps per connection) and is chosen when the customer cannot meet Google at a colocation facility. Option B is wrong because Direct Peering is a BGP peering relationship for exchanging traffic, not a dedicated private connection with an SLA, and it does not provide the 99.99% availability guarantee. Option D is wrong because HA VPN is an IPsec VPN over the public internet with a 99.99% SLA but maximum throughput of 3 Gbps per tunnel, far below the 100 Gbps requirement.

37
MCQhard

A healthcare company runs a critical patient portal on Google Kubernetes Engine. The security team requires that all container images be scanned for vulnerabilities before deployment, that only images from a trusted registry be admitted to the cluster, and that any attempt to deploy an untrusted image be blocked and logged. Which Google Cloud feature should the cloud architect implement to enforce these admission requirements?

A.GKE network policies that restrict egress from the cluster to only the trusted registry, preventing pods from pulling images from unapproved sources.
B.GKE Sandbox with gVisor runtime enabled on all node pools to isolate containers and reduce the impact of vulnerable images.
C.Binary Authorization with a policy that requires attestations from a trusted vulnerability scanner and allows only images from the approved registry, integrated with GKE admission control.
D.Artifact Registry vulnerability scanning with automatic scanning enabled on push for all repositories.
AnswerC

Binary Authorization enforces deploy-time policies on GKE by requiring cryptographic attestations that prove an image passed the required scanning step and came from an approved registry. When a deployment violates the policy, admission is denied and the attempt is logged. This directly meets the need to block and record untrusted image deployments.

Why this answer

Binary Authorization is the GKE-native admission control that enforces policies at deploy time using attestations. A policy can require that images be attested by a trusted scanner and originate from an approved registry, and violations are denied and logged. Vulnerability scanning alone detects but does not block, network policies do not govern kubelet image pulls, and sandboxing hardens runtime without validating provenance.

Exam trap

The trap here is assuming that enabling vulnerability scanning on a registry is enough to prevent vulnerable images from running, when scanning only reports findings and requires an admission controller to enforce them.

38
MCQeasy

A company uses Cloud Logging to capture application logs. They need to alert when the number of errors exceeds 100 in a 5-minute window. Which type of alert should they create?

A.Notification channel with email integration
B.Cloud Logging sink to a Pub/Sub topic
C.Log-based metric with an alerting policy
D.SLO alerting policy
AnswerC

A log-based metric counts matching log entries, such as errors, and an alerting policy on that metric triggers when the count exceeds 100 within the 5-minute window. This satisfies the threshold condition on error volume.

Why this answer

To alert on the count of error log entries exceeding 100 in a 5-minute window, the engineer must first create a log-based metric that counts matching log entries, then attach an alerting policy with a threshold condition on that metric. Cloud Logging alone cannot alert on log content; it must be converted into a metric that Cloud Monitoring can evaluate. This is the standard pattern for log-driven alerting in Google Cloud.

Exam trap

PCA often tests whether candidates know that Cloud Logging cannot alert directly on log content and that a log-based metric plus an alerting policy is required.

How to eliminate wrong answers

Option A is wrong because a notification channel only defines where alerts are sent; it does not define the condition or the metric, so it cannot trigger on error counts by itself. Option B is wrong because a log sink to Pub/Sub routes log entries for downstream processing but does not evaluate thresholds or generate alerts. Option D is wrong because an SLO alerting policy is based on service-level objectives (error budgets, burn rates) and is not designed to count raw error log entries in a 5-minute window.

39
MCQhard

A healthcare company runs a critical application on Google Kubernetes Engine (GKE) that processes patient data. The compliance team requires that all container images be scanned for vulnerabilities before deployment, and that only images from a trusted registry be allowed. The security team wants to enforce this policy across all clusters in the organization. They also need to audit any attempts to deploy untrusted images. Which combination of Google Cloud services should they use?

A.Enable GKE Sandbox on all nodes and use Anthos Config Management to apply a policy that restricts image registries.
B.Use Container Analysis to scan images and configure a GKE admission controller to reject images with vulnerabilities.
C.Use Artifact Registry with vulnerability scanning enabled and configure IAM policies to restrict which projects can pull images.
D.Use Binary Authorization with a policy that requires attestations from a trusted authority, and enable audit logging for GKE.
AnswerD

Binary Authorization enforces deploy-time policies on GKE by requiring attestations that prove an image was built by a trusted builder and scanned for vulnerabilities. By configuring a policy that requires attestations from a trusted authority, the company ensures only compliant images are deployed. Enabling audit logging captures attempts to deploy non-compliant images, satisfying the audit requirement. This combination directly addresses both enforcement and auditing across all clusters in the organization.

Why this answer

Binary Authorization is the Google Cloud service designed to enforce deploy-time policies on GKE by requiring attestations. By setting a policy that requires attestations from a trusted authority, the company ensures that only images that have been built and scanned by trusted processes are admitted. Enabling audit logging provides a record of all deployment attempts, including those that violate the policy.

This satisfies both the enforcement and auditing requirements across all clusters in the organization.

Exam trap

The trap here is confusing vulnerability scanning with policy enforcement; scanning alone does not prevent deployment of vulnerable images.

40
MCQhard

An application running on GKE Autopilot is experiencing intermittent failures due to resource limits. The team wants to ensure that the application always has enough CPU and memory without manual node management. What should they do?

A.Use horizontal pod autoscaling only
B.Increase the resource requests and limits in the pod specification
C.Create a new node pool with larger machine types
D.Switch to GKE Standard and manage node pools manually
AnswerB

Raising resource requests and limits in the pod specification directly addresses the intermittent failures by guaranteeing the scheduler reserves sufficient CPU and memory for each pod. On GKE Autopilot, nodes are provisioned automatically to match those requests, so adequate values remove throttling and OOM evictions without any manual node management.

Why this answer

In GKE Autopilot, nodes are fully managed by Google, so the team cannot create or resize node pools. The correct lever is to set appropriate CPU and memory requests (and limits) in the pod specification so the scheduler and Autopilot's autoscaler provision nodes with sufficient capacity. Autopilot uses the requests to bin-pack pods and to decide when to add nodes, so undersized requests cause intermittent failures under load.

Exam trap

The trap is treating GKE Autopilot like GKE Standard — candidates pick 'create a larger node pool' or 'switch to Standard' because that is the Standard-mode answer, forgetting that Autopilot hides node management and the only correct lever is pod resource requests/limits.

How to eliminate wrong answers

Option A is wrong because horizontal pod autoscaling only adjusts the number of pod replicas based on metrics; it does not address per-pod resource limits and can even worsen failures if pods are OOM-killed due to insufficient memory requests. Option C is wrong because GKE Autopilot does not expose node pools for manual creation or machine-type selection — that is a GKE Standard capability. Option D is wrong because switching to GKE Standard and managing node pools manually abandons the Autopilot operational model the team is using and is unnecessary to solve the problem.

41
MCQeasy

An engineer needs to create a custom dashboard in Cloud Monitoring to track the 99th percentile latency of their application over the last 7 days. Which type of metric should they use?

A.Distribution metric
B.Delta metric
C.Cumulative metric
D.Gauge metric
AnswerA

Distribution metrics record a histogram of values across a time window, letting Cloud Monitoring compute percentiles such as p99 directly. This satisfies the requirement to track 99th percentile latency over seven days, which a gauge or counter cannot represent.

Why this answer

Distribution metrics capture a histogram of values across a population of samples, preserving the full value distribution rather than collapsing it to a single number. This is essential for computing percentile aggregations like p99, since percentiles require the underlying histogram buckets to be calculated. Cloud Monitoring's distribution metrics (e.g., from OpenCensus/OpenTelemetry or load balancer latency) support aligner/aggregation functions such as percentile, count, and mean over a time window.

Exam trap

PCA often tests the confusion between metric kinds — candidates pick gauge because they think 'latency is a single value,' forgetting that percentiles require the distribution kind to be computed at all.

How to eliminate wrong answers

Option B is wrong because delta metrics record the change in a cumulative counter between samples, which is useful for rate calculations but cannot produce percentile statistics. Option C is wrong because cumulative metrics monotonically increase from a start time and are designed for rate/derivative operations, not distribution analysis. Option D is wrong because gauge metrics represent a single instantaneous value (like current CPU usage) and have no distribution to compute percentiles from.

42
MCQhard

A financial services company runs a critical PostgreSQL database on Cloud SQL. They need to ensure automatic failover to a replica in another zone within the same region with minimal data loss. What configuration should they choose?

A.Use Database Migration Service to replicate to a second Cloud SQL instance
B.Enable point-in-time recovery (PITR) and increase backup retention
C.Create a cross-region read replica and manually promote it on failure
D.Configure a Cloud SQL HA instance with a failover replica in a different zone
AnswerD

A Cloud SQL HA instance maintains a standby in a different zone with synchronous replication, so failover is automatic and data loss is minimal. This satisfies the stem's requirement for cross-zone automatic failover within the same region.

Why this answer

Cloud SQL High Availability (HA) provisions a standby instance in a different zone within the same region and performs automatic failover with synchronous replication, giving near-zero RPO and minimal downtime. This directly satisfies the requirement for automatic cross-zone failover with minimal data loss.

Exam trap

The trap is conflating 'replica' with 'HA failover' — candidates pick a cross-region read replica thinking it provides automatic failover, but read replicas are asynchronous and require manual promotion, unlike the synchronous HA standby.

How to eliminate wrong answers

Option A is wrong because Database Migration Service is for one-time or continuous migration into Cloud SQL, not for providing an automatic failover target. Option B is wrong because PITR and backup retention only enable point-in-time restore after an incident — recovery is manual and can take many minutes, with data loss up to the last WAL/binlog. Option C is wrong because a cross-region read replica requires manual promotion and is asynchronous, so it does not meet the 'automatic failover within the same region with minimal data loss' requirement.

43
MCQmedium

A company uses Cloud SQL for MySQL for its transactional database. They need to ensure automatic failover in case of a zonal outage with minimal data loss. What configuration should they use?

Answer options not yet available.

Why this answer

Cloud SQL High Availability (HA) configuration creates a standby instance in a different zone within the same region. If the primary fails, it automatically fails over to the standby, minimizing downtime. Backup and PITR help with data loss but do not provide automatic failover.

44
MCQeasy

A company needs to retain object versions in Cloud Storage for 90 days to protect against accidental deletion or modification. After 90 days, versions should be deleted. What feature should they enable?

A.Object versioning only
B.Retention policy
C.Object holds
D.Object lifecycle management with a rule to delete versions after 90 days
AnswerD

Object lifecycle management applies age-based rules to noncurrent object versions, automatically deleting them once they pass 90 days. This satisfies the stem's requirement to retain versions for 90 days and then remove them, which a retention policy alone cannot do.

Why this answer

Object Lifecycle Management lets you define rules to automatically delete object versions after a specified age (90 days). Combined with Object Versioning enabled on the bucket, this retains noncurrent versions for 90 days and then deletes them, protecting against accidental deletion or modification while controlling storage costs.

Exam trap

The trap is confusing Retention Policy (which locks objects and prevents deletion) with Lifecycle Management (which deletes versions) — candidates may pick Retention Policy thinking it deletes after the period, but it does the opposite.

How to eliminate wrong answers

Option A is wrong because Object Versioning alone retains versions indefinitely, causing unbounded storage growth and cost. Option B is wrong because a Retention Policy locks objects for a period but does not delete them after the period; it also prevents deletion, which is the opposite of the requirement to delete after 90 days. Option C is wrong because Object Holds (temporary or event-based) prevent deletion until released, but do not automatically delete after 90 days.

45
MCQhard

A financial services company runs a latency-sensitive trading application on GKE. The platform team must guarantee that the application can be recovered within a 15-minute recovery time objective (RTO) and a 5-minute recovery point objective (RPO) after a regional failure. They use a multi-region Cloud Storage bucket for configuration and a regional GKE cluster. Which additional design element is required to meet both objectives?

A.Enable GKE cluster autoscaling and configure a horizontal pod autoscaler for the trading pods.
B.Configure the existing regional GKE cluster with a node pool in each of the three zones of its region and enable pod anti-affinity.
C.Use Anthos Config Management to sync configurations to the existing cluster and enable Binary Authorization for the trading images.
D.Deploy a second regional GKE cluster in another region and use a multi-region Cloud Storage bucket for shared configuration, with a documented failover runbook.
AnswerD

A standby regional GKE cluster in a different region provides compute capacity to restart workloads when the primary region fails. Because configuration is stored in a multi-region Cloud Storage bucket, it remains accessible during the failover, supporting a short RPO. A documented runbook ensures the team can perform failover within the 15-minute RTO.

Why this answer

Meeting a 15-minute RTO and 5-minute RPO after a regional failure requires compute capacity and configuration data available outside the failed region. A standby GKE cluster in another region supplies the compute, while a multi-region Cloud Storage bucket keeps configuration accessible. A rehearsed failover runbook ensures the team can switch over quickly enough to satisfy the RTO.

Exam trap

The trap here is treating multi-zone node pools within a single region as sufficient for a regional failure, when they only protect against zone-level outages.

46
MCQmedium

A company has a Cloud SQL for PostgreSQL instance in a single zone. To achieve high availability, they want to ensure automatic failover with zero data loss and minimal downtime. Which configuration should they use?

A.Deploy a read replica in the same zone and enable automatic failover
B.Enable automatic backups and point-in-time recovery
C.Add a cross-region read replica and configure failover manually
D.Configure a Cloud SQL regional instance with a failover replica in a different zone
AnswerD

A Cloud SQL regional instance maintains a synchronous standby replica in a different zone within the same region. Synchronous replication guarantees zero data loss (RPO of zero), while automatic failover to the standby delivers minimal downtime, satisfying both the high availability and no-data-loss constraints in the stem.

Why this answer

A Cloud SQL regional instance provisions a standby replica in a different zone within the same region and performs automatic failover with synchronous replication, ensuring zero data loss (RPO=0) and minimal downtime. This is the standard HA configuration for Cloud SQL for PostgreSQL.

Exam trap

PCA often tests the misconception that read replicas provide HA failover; in Cloud SQL, only a regional instance with a standby replica delivers automatic zero-data-loss failover.

How to eliminate wrong answers

Option A is wrong because read replicas use asynchronous replication and cannot be promoted automatically as an HA failover target — they are for read scaling, not HA. Option B is wrong because backups and PITR address data recovery after corruption or deletion, not automatic failover; they involve restore downtime and potential data loss up to the last backup/binlog. Option C is wrong because cross-region read replicas are for disaster recovery and read offload, and failover (promotion) is a manual operation with asynchronous replication, so it does not meet zero-data-loss or automatic failover requirements.

47
MCQhard

An e-commerce platform uses Cloud Spanner in a multi-region configuration. They want to achieve the highest possible availability SLA. Which deployment configuration should they choose?

Answer options not yet available.

Why this answer

Cloud Spanner offers a 99.999% SLA for multi-region configurations. To achieve this, you must use a multi-region instance (e.g., nam3, eur3) that replicates data across at least three regions. A single-region configuration only offers 99.99% SLA.

48
MCQmedium

Your team uses Cloud SQL for PostgreSQL for an e-commerce application. You want to perform point-in-time recovery (PITR) to recover from a logical error that occurred 10 minutes ago. Which prerequisites are required?

A.Automated backups must be enabled, and the instance must be using the InnoDB storage engine
B.Automated backups and binary logging must be enabled
C.Point-in-time recovery is not supported for Cloud SQL PostgreSQL
D.Automated backups must be enabled, and write-ahead logging (WAL) must be active
AnswerD

Automated backups and write-ahead logging are the two prerequisites Cloud SQL for PostgreSQL requires for point-in-time recovery. WAL archiving captures continuous transaction logs, while automated backups supply the base snapshot; together they let you restore to any moment within the retention window, satisfying the 10-minute recovery target.

Why this answer

For Cloud SQL for PostgreSQL, point-in-time recovery (PITR) requires that automated backups are enabled and that write-ahead logging (WAL) is active. WAL archiving captures all changes, allowing recovery to any point within the retention period. Automated backups provide the base backup from which WAL is replayed.

Exam trap

The trap is confusing MySQL terminology (binary logging, InnoDB) with PostgreSQL (WAL), or assuming PITR is not supported for PostgreSQL.

How to eliminate wrong answers

Option A is wrong because InnoDB is a MySQL storage engine, not PostgreSQL. Option B is wrong because binary logging is a MySQL concept; PostgreSQL uses WAL. Option C is wrong because PITR is supported for Cloud SQL PostgreSQL when prerequisites are met.

49
MCQeasy

You want to create a log-based alert in Cloud Logging that triggers when a specific error message appears in application logs. What is the first step?

A.Create a logs-based metric that filters for the error message
B.Configure a Pub/Sub notification channel for alerts
C.Create a log sink to export logs to Cloud Storage
D.Set up an alerting policy directly on the log entries without a metric
AnswerA

A logs-based metric counts matching log entries, and an alerting policy can only fire against such a metric. Creating it first, filtered on the specific error message, satisfies the trigger requirement before any notification channel or alert policy is configured.

Why this answer

To create a log-based alert, you first define a logs-based metric that counts occurrences of the error pattern. Then you create an alerting policy that monitors this metric and triggers when the count exceeds a threshold. Notifications are configured in the alerting policy, not the metric.

50
MCQmedium

Your organization runs a global e-commerce platform on Google Kubernetes Engine (GKE). The security team requires that all container images deployed to the cluster are scanned for vulnerabilities and that deployments are blocked if critical vulnerabilities are found. They also want to minimize operational overhead. What should you do?

A.Configure a Kubernetes admission controller that calls the Container Analysis API to check for vulnerabilities and rejects pods with critical findings.
B.Enable Pod Security Policies to restrict images to those from trusted registries, and rely on registry scanning to prevent vulnerable images.
C.Use Cloud Build to scan images with Container Analysis, and manually review the scan results before approving each deployment.
D.Enable Binary Authorization in the cluster and configure a policy that requires attestations from a vulnerability scanner before deployment.
AnswerD

Binary Authorization enforces deploy-time security controls by verifying attestations. You can integrate a vulnerability scanner (e.g., Container Analysis) to create attestations only for images that pass scanning. This blocks non-compliant images and reduces manual effort, aligning with the requirement to block critical vulnerabilities with minimal overhead.

Why this answer

Binary Authorization is a Google Cloud service that enforces deploy-time policies by requiring attestations. By integrating with Container Analysis, you can automatically attest only images that pass vulnerability scanning, and the policy blocks images without attestations. This automates enforcement and reduces manual review, satisfying both security and operational efficiency.

Exam trap

The trap here is assuming that vulnerability scanning alone can block deployments, when in fact scanning only reports findings and requires an enforcement mechanism like Binary Authorization to prevent deployment.

51
Multi-Selecteasy

An engineer needs to troubleshoot a production issue on a Compute Engine instance. They suspect the instance is running out of memory. Which THREE actions should they take to diagnose the problem? (Choose THREE.)

Select 3 answers
A.SSH into the instance and run 'free -m' to check memory usage
B.Check Cloud Logging for OOM (out-of-memory) kernel messages
C.Increase the instance's memory by changing the machine type
D.Create a snapshot of the boot disk
E.View the instance's memory utilization metric in Cloud Monitoring
AnswersA, B, E

Running 'free -m' over SSH reads the instance's actual memory counters, showing total, used, free and swap usage plus buffer/cache. This directly confirms or rules out memory exhaustion on the guest OS, which external metrics alone cannot attribute to specific processes.

Why this answer

Option A is correct because running 'free -m' over SSH directly reports the instance's current RAM and swap usage in megabytes, immediately confirming whether memory is exhausted. Option B is correct because when the Linux kernel OOM killer terminates processes, it logs messages such as 'Out of memory: Killed process' to the kernel log, which is captured by the Cloud Logging agent and searchable in Logs Explorer. Option E is correct because the Cloud Monitoring agent (or Ops Agent) publishes the 'memory/utilization' metric, letting the engineer review historical memory trends and correlate spikes with the incident.

Option C is not a diagnostic action but a remediation that changes the machine type and requires a stop/start, so it does not help identify the cause. Option D is also not diagnostic; creating a boot disk snapshot only preserves disk state and does not reveal memory consumption.

52
MCQeasy

An engineer needs to list all Compute Engine instances in a project using the command line. Which gcloud command should they use?

A.gcloud compute instances describe
B.gcloud compute instances list
C.gcloud compute instance-groups list
D.gcloud compute machine-types list
AnswerB

The gcloud compute instances list command queries the Compute Engine API and returns every instance in the active project, with optional filtering by zone or name. It directly satisfies the stem's requirement to enumerate all Compute Engine instances from the command line.

Why this answer

The correct command to list Compute Engine instances is 'gcloud compute instances list'. The other options are incorrect: 'gcloud compute machine-types list' lists machine types, 'gcloud compute instance-groups list' lists instance groups, and 'gcloud compute instances describe' describes a specific instance.

53
MCQeasy

A team is adopting a DevOps model and wants to reduce the risk of configuration drift between environments. They deploy the same application to development, staging, and production projects on Google Cloud. Which practice should they adopt to ensure consistent, repeatable deployments across all environments?

A.Manually apply changes in each project and record them in a shared spreadsheet for audit purposes.
B.Use Cloud Console to configure each project and rely on Cloud Asset Inventory to detect differences after deployment.
C.Grant all engineers Owner role on each project so they can quickly recreate resources when drift is detected.
D.Define infrastructure as code using Terraform with separate variable files per environment, and store the configuration in a version control system.
AnswerD

Infrastructure as code with Terraform makes deployments declarative and repeatable. Separate variable files let the same modules target development, staging, and production with environment-specific values, while version control provides history, review, and rollback. This approach minimizes drift because the desired state is defined once and applied consistently.

Why this answer

Repeatable deployments across environments require a declarative, versioned definition of infrastructure. Terraform with per-environment variable files lets one set of modules produce consistent resources in development, staging, and production, while version control records changes and enables review and rollback. This prevents configuration drift far more effectively than manual changes or post-hoc detection.

Exam trap

The trap here is confusing drift detection with drift prevention, assuming that inventory or audit tools alone can keep environments consistent.

54
MCQmedium

Your organization runs a production Cloud SQL for PostgreSQL instance. You need to ensure that if the primary zone fails, the database automatically fails over to a standby with no data loss. Which configuration should you use?

A.Enable point-in-time recovery (PITR)
B.Configure a cross-region replica
C.Deploy a regional Cloud SQL instance with high availability
D.Create a read replica and promote it on failure
AnswerC

A regional Cloud SQL instance with high availability maintains a standby in a separate zone and performs automatic failover with synchronous replication, ensuring no data loss when the primary zone fails. Zonal instances lack this standby capability.

Why this answer

A regional Cloud SQL instance with high availability provisions a standby in a different zone within the same region and performs automatic failover via a regional persistent disk, with synchronous replication ensuring zero data loss (RPO ≈ 0). This is the only option that combines automatic failover with no data loss for a single-region production database.

Exam trap

PCA often tests the distinction between HA (automatic, synchronous, zero data loss, same region) and replicas/PITR (asynchronous or manual, potential data loss), so candidates who pick 'read replica and promote' or 'PITR' for failover fall into the trap.

How to eliminate wrong answers

Option A is wrong because point-in-time recovery only lets you restore to an earlier timestamp after the fact; it does not provide automatic failover and typically incurs data loss up to the last recoverable point. Option B is wrong because a cross-region replica is asynchronous, so failover would lose recent transactions and it is not an automatic HA mechanism. Option D is wrong because a read replica is asynchronous and promotion is a manual, lossy operation, not automatic failover.

55
MCQmedium

Your company runs a stateful application on GKE that stores data in persistent volumes backed by Compute Engine persistent disks. You need to back up the application data and the Kubernetes resource configurations (deployments, services, etc.) for disaster recovery. Which tool should you use?

A.Velero
B.gcloud container clusters create --async
C.Cloud SQL for MySQL
D.Cloud Storage with object versioning and lifecycle policies
AnswerA

Velero backs up both Kubernetes resource configurations and persistent volume data, integrating with Compute Engine persistent disk snapshots. That combination satisfies the disaster recovery requirement to restore deployments, services and application data together on GKE.

Why this answer

Velero is the standard open-source tool for backing up and restoring Kubernetes cluster resources and persistent volumes. It can snapshot Compute Engine persistent disks and also back up Kubernetes objects like deployments and services, which matches the requirement to protect both application data and resource configurations.

Exam trap

PCA often tests the misconception that generic storage versioning or database backups cover Kubernetes disaster recovery, when the exam expects a Kubernetes-aware tool like Velero for both resources and persistent volumes.

How to eliminate wrong answers

Option B is wrong because gcloud container clusters create --async creates a new cluster asynchronously and has nothing to do with backup or disaster recovery. Option C is wrong because Cloud SQL for MySQL is a managed relational database service and cannot back up GKE persistent volumes or Kubernetes manifests. Option D is wrong because Cloud Storage with object versioning and lifecycle policies stores objects but does not natively back up Kubernetes resources or persistent disk snapshots in a restorable, cluster-aware way.

56
Multi-Selectmedium

A company wants to implement a disaster recovery (DR) strategy for their Cloud SQL for MySQL databases. They need to be able to recover to a specific point in time (within seconds) in case of accidental data deletion. Which TWO actions should they take? (Choose TWO.)

Select 2 answers
A.Enable binary logging (binlog)
B.Configure a failover replica in another zone
C.Create a cross-region read replica
D.Enable automated backups
E.Export the database daily to Cloud Storage
AnswersA, D

Binary logging captures every committed transaction, enabling Cloud SQL point-in-time recovery to a chosen timestamp within seconds of accidental deletion. This satisfies the seconds-level recovery objective, which automated backups alone cannot meet because they restore only to fixed snapshot times.

Why this answer

Option A is correct because Cloud SQL for MySQL point-in-time recovery (PITR) relies on binary logging (binlog) to capture all data changes, allowing recovery to a specific moment within seconds. Option D is correct because automated backups provide the base backup that, combined with binlog, enables PITR; without automated backups enabled, PITR cannot function. Option B is incorrect because a failover replica provides high availability within a region, not point-in-time recovery from accidental deletion.

Option C is incorrect because a cross-region read replica is for read scaling and regional DR, not for recovering to a specific point in time. Option E is incorrect because daily exports to Cloud Storage are manual/periodic snapshots and cannot achieve second-level point-in-time recovery.

Exam trap

PCA often tests the confusion between high availability (failover replicas) and point-in-time recovery, tricking candidates into selecting HA features that replicate — rather than protect against — accidental data deletion.

57
MCQhard

Your company uses Cloud Spanner in a multi-region configuration to achieve 99.999% availability. You need to understand the impact of a regional failure on read and write availability. Which statement is correct?

A.Both reads and writes are fully available as long as at least one region remains healthy
B.Reads and writes remain fully available because Cloud Spanner uses synchronous replication across all regions
C.Writes are unavailable if the region containing the leader replica fails, but reads remain available
D.Writes are always available, but reads may be unavailable if the region with the closest replica fails
AnswerC

Cloud Spanner's leader replica handles all writes, so a regional failure affecting that leader halts writes until a new leader is elected. Reads, however, are served by any replica, so they continue from surviving regions, satisfying the stem's read-versus-write availability distinction.

Why this answer

Cloud Spanner multi-region configurations use a leader replica in one region and read-only replicas in others. Writes must go through the leader (Paxos quorum), so if the region hosting the leader fails, writes are unavailable until a new leader is elected — typically within seconds to minutes depending on the config. Reads, however, can be served from any healthy replica (strong or stale reads), so read availability is preserved as long as at least one replica region survives.

Exam trap

PCA often tests the misconception that synchronous replication equals write availability everywhere — candidates forget that only the leader accepts writes and pick 'both reads and writes fully available'.

How to eliminate wrong answers

Option A is wrong because it claims writes remain fully available during any single-region failure, ignoring the leader-election dependency — writes stall during leader failover. Option B is wrong because synchronous replication does not mean every region can accept writes; only the leader does, and synchronous replication is precisely why writes pause during leader failover. Option D is wrong because it inverts the model: reads are the more available operation (any replica can serve them), while writes are the constrained operation tied to the leader.

58
Multi-Selecthard

A financial services company runs a latency-sensitive trading application on Compute Engine. The operations team needs to detect performance regressions and correlate them with recent deployments without instrumenting application code. They want to use Cloud Monitoring and Cloud Logging features that work automatically for Compute Engine VMs. (Choose two.)

Select 2 answers
A.Install the Cloud Profiler agent on each VM to capture CPU and memory profiles of the application.
B.Create log-based metrics in Cloud Logging from VM logs and chart them alongside metrics in Cloud Monitoring dashboards.
C.Enable Data Access audit logs for Compute Engine and analyze them to identify slow API calls.
D.Use the Ops Agent to collect host metrics such as CPU, memory, and disk I/O, and to send application logs to Cloud Logging.
E.Enable Cloud Trace auto-instrumentation by setting the GOOGLE_CLOUD_TRACE environment variable on each VM.
AnswersB, D

Log-based metrics turn log entries into time-series data that Cloud Monitoring can chart and alert on. Combined with the Ops Agent's log collection, this lets the team correlate application log patterns with performance metrics over time without modifying application code, which directly supports detecting regressions relative to deployments.

Why this answer

The Ops Agent provides automatic host metrics and log collection for Compute Engine, and log-based metrics let those logs be charted and alerted in Cloud Monitoring. Together they deliver the telemetry to detect performance regressions and correlate them with deployments without touching application code. Tracing, profiling, and audit logs either require code changes or capture different data than runtime performance.

Exam trap

The trap here is assuming that enabling tracing or profiling is automatic on Compute Engine, when those require application-level instrumentation and do not satisfy the no-code-change constraint.

59
MCQeasy

You need to automatically roll back a GKE deployment if a new version causes a spike in 5xx errors. The deployment uses a canary strategy with Istio traffic splitting. What should you do?

A.Use Cloud Monitoring to watch the canary's error rate and trigger a Cloud Function that updates the Istio VirtualService to route all traffic back to the stable version.
B.Set the canary's traffic weight to 0 in the Istio VirtualService if errors exceed threshold using a Kubernetes Job.
C.Use GKE's built-in auto-repair feature to replace unhealthy pods.
D.Configure an Istio VirtualService with a retry policy that automatically redirects traffic on errors.
AnswerA

Cloud Monitoring detects the canary's elevated 5xx rate, and the triggered Cloud Function rewrites the Istio VirtualService weights, shifting all traffic back to the stable version. This satisfies the automatic rollback requirement without redeploying, since Istio controls routing independently of the GKE workload.

Why this answer

The correct pattern is to monitor the canary's error rate with Cloud Monitoring (using Istio's telemetry metrics like istio_requests_total filtered by response_code=5xx and destination_version=canary), then use an alerting policy to trigger a Cloud Function (or Cloud Run) that patches the Istio VirtualService to shift 100% of traffic back to the stable version. This closes the loop between observability and traffic control, which is exactly what automated canary rollback requires.

Exam trap

PCA often tests the misconception that Kubernetes auto-repair or Istio retries provide application-level rollback — candidates pick them because they sound like resilience features, but neither changes traffic routing based on error rates.

How to eliminate wrong answers

Option B is wrong because a Kubernetes Job is not event-driven — it runs to completion and cannot react to a live error-rate spike without an external trigger, so it cannot perform timely rollback. Option C is wrong because GKE auto-repair only restarts/replaces pods that fail health checks; it does not detect application-level 5xx spikes or change traffic routing, so a canary returning 500s while passing liveness probes would never be rolled back. Option D is wrong because Istio retry policies retry failed requests to the same destination; they do not redirect traffic to a different version and can amplify load during an outage rather than rolling back.

60
MCQmedium

A retail company runs an e-commerce platform on GKE. The SRE team wants to measure the error budget for a service level objective (SLO) defined as the proportion of requests served with HTTP 2xx or 3xx status over a 28-day window. They need a monitoring configuration that computes the burn rate and alerts when the budget is being consumed too quickly, while avoiding noisy alerts during brief spikes. What should they do?

A.Create an alerting policy on a log-based metric that counts 5xx responses and alert when the count exceeds a static threshold.
B.Use Cloud Monitoring SLO monitoring to define the request-based SLO, then create a multi-window, multi-burn-rate alerting policy on that SLO.
C.Export request metrics to BigQuery and run a scheduled query that emails the team when the 28-day success ratio drops below the target.
D.Create a Cloud Monitoring uptime check on the service endpoint and alert when two consecutive checks fail.
AnswerB

Cloud Monitoring SLO monitoring supports request-based SLOs and calculates error budgets and burn rates. A multi-window, multi-burn-rate alerting policy fires on fast and slow burn conditions, which reduces noise from brief spikes while still catching sustained budget consumption. This matches the requirement to alert on rapid budget depletion with fewer false positives.

Why this answer

Cloud Monitoring SLO monitoring natively models request-based SLOs, tracks error budgets, and supports burn-rate alerting. Multi-window, multi-burn-rate policies combine a short window for fast detection with a longer window for confirmation, which filters out brief spikes. This is the supported, least-effort way to alert on rapid budget consumption for a 28-day SLO.

Exam trap

The trap here is treating a static threshold on error counts or a simple uptime check as equivalent to burn-rate alerting, when only SLO-based monitoring accounts for the request ratio over the SLO window.

61
MCQhard

Your team is following an incident management process. After resolving a major incident, you are tasked with conducting a postmortem. What is the PRIMARY goal of the postmortem process in Google Cloud's recommended approach?

A.Understand the root cause and implement changes to prevent recurrence
B.Document the incident timeline and communicate it to stakeholders
C.Calculate the financial impact and bill the responsible team
D.Identify the individual responsible for the incident and take corrective action
AnswerA

Google Cloud's postmortem process is blameless and focuses on identifying the underlying root cause of the incident, then implementing corrective actions so the same failure cannot recur. This satisfies the primary goal of preventing recurrence rather than assigning fault.

Why this answer

Google's recommended postmortem process, derived from SRE practices, is blameless and focused on learning: the primary goal is to understand the root cause(s) of the incident and implement systemic changes to prevent recurrence. It emphasizes identifying contributing factors across technology, process, and human dimensions rather than assigning fault. Documentation and communication are outputs, not the primary goal, and financial or punitive outcomes are explicitly excluded.

Exam trap

PCA often tests the blameless principle — candidates pick the option about identifying the responsible individual, but Google's postmortem explicitly rejects blame in favor of systemic learning.

How to eliminate wrong answers

Option B is wrong because documenting the timeline and communicating to stakeholders is a necessary output of the postmortem, but it is not the primary goal — the goal is learning and prevention, not reporting. Option C is wrong because calculating financial impact and billing a team contradicts Google's blameless culture and is not part of the postmortem process; cost analysis may occur separately but is not the objective. Option D is wrong because identifying an individual to blame and taking corrective action is explicitly antithetical to Google's blameless postmortem philosophy, which holds that blaming individuals discourages transparency and hides systemic issues.

62
MCQeasy

You need to create a Cloud Logging sink that exports logs to a BigQuery dataset for long-term analysis. Which destination type should you specify?

A.Cloud Storage
B.BigQuery
C.Pub/Sub
D.Custom HTTP endpoint
AnswerB

Cloud Logging sinks route log entries to supported destinations, and BigQuery is a native sink destination that stores exported logs in datasets for SQL analysis and long-term retention. Specifying BigQuery as the destination type satisfies the requirement to export logs into a BigQuery dataset.

Why this answer

Cloud Logging sinks support three destination types: Cloud Storage, BigQuery, and Pub/Sub (plus custom destinations via Pub/Sub or Logging API). BigQuery is the correct choice here because the requirement is long-term analysis of log data, and BigQuery provides a fully managed, serverless data warehouse with SQL querying, partitioning, and clustering capabilities ideal for analytical workloads on exported logs.

Exam trap

PCA often tests the distinction between log routing destinations by matching the destination to the use case — candidates incorrectly pick Pub/Sub for 'analysis' when the question implies batch SQL analytics, which is BigQuery's domain.

How to eliminate wrong answers

Option A is wrong because Cloud Storage is designed for object storage and archival, not for running analytical SQL queries over log data — it is best suited for cold storage or log retention rather than analysis. Option C is wrong because Pub/Sub is a messaging service used to stream logs to downstream consumers in real time, not a storage or analytics destination for long-term analysis. Option D is wrong because a custom HTTP endpoint is not a native Logging sink destination type; custom destinations require routing through Pub/Sub or the Logging API.

63
MCQeasy

Your organization wants to use Cloud SQL for a MySQL database with automatic failover in the event of a zone outage. Which configuration should you choose?

A.Set up Cloud SQL with external replication to a VM in another zone
B.Create a Cloud SQL instance with a cross-region read replica
C.Create a single-zone Cloud SQL instance with automatic backups enabled
D.Create a regional Cloud SQL instance (high availability) with a primary and standby zone
AnswerD

A regional instance maintains a standby in a different zone within the same region, with synchronous replication and automatic failover. This satisfies the zone-outage requirement, whereas a zonal instance has no standby and cannot fail over automatically.

Why this answer

A regional Cloud SQL instance (high availability) provisions a primary and a standby instance in two zones within the same region, with automatic failover to the standby if the primary zone fails. This is the native Cloud SQL HA configuration for zone-outage resilience. Cross-region read replicas and external replication are for read scaling or DR, not automatic same-region failover.

Exam trap

PCA often tests the difference between HA (regional instance with standby for automatic failover) and read replicas (for scaling/DR) — candidates pick cross-region replicas thinking they provide automatic failover.

How to eliminate wrong answers

Option A is wrong because external replication to a VM in another zone is a manual, self-managed setup that does not provide Cloud SQL's automatic failover. Option B is wrong because a cross-region read replica is for read scaling and disaster recovery across regions, not automatic failover within a region. Option C is wrong because a single-zone instance with automatic backups provides data recovery via restore, but no automatic failover during a zone outage — the instance goes down until the zone recovers or you manually restore.

64
MCQmedium

An organization needs to run a stateful application on Google Kubernetes Engine (GKE) where the nodes are fully managed by Google and the application workload SLAs are guaranteed. They want to minimize operational overhead. Which GKE mode should they use?

A.GKE Standard with Cluster Autoscaler
B.GKE Standard with node auto-provisioning
C.GKE Standard with sole-tenant nodes
D.GKE Autopilot
AnswerD

GKE Autopilot provisions and manages the node infrastructure itself, including scaling, patching and node pool configuration, while enforcing workload resource requests and SLA-backed reliability. This removes node-level operational overhead, satisfying the requirement for fully Google-managed nodes with guaranteed workload SLAs.

Why this answer

GKE Autopilot manages the entire node infrastructure including node provisioning, scaling, and maintenance. It provides workload-level SLAs (e.g., 99.95% for pods). Standard mode requires the user to manage node pools.

65
MCQmedium

A team is migrating a monolithic application to microservices on GKE. They want to gradually shift users to the new microservices version while keeping the old monolithic version running. They need to route a small percentage of users based on a cookie. Which traffic management approach should they use?

A.Use Istio VirtualService with match rules based on cookie and weighted destinations
B.Use Kubernetes Services with multiple Deployments and manual scaling
C.Configure an HTTP(S) load balancer with URL maps and backend services
D.Deploy two separate GKE clusters and use DNS-based traffic splitting
AnswerA

Istio VirtualService supports match conditions on request headers such as cookies, combined with weighted destination routing, enabling a defined percentage of cookie-matched users to reach the microservices version while the remainder continue to the monolith. This satisfies the gradual, cookie-based traffic shift.

Why this answer

Istio traffic management allows fine-grained routing based on HTTP headers, cookies, or other attributes. It supports traffic splitting and canary deployments with precise percentage control.

66
MCQhard

An e-commerce company runs its order-processing service on Cloud Run. During flash sales, the service experiences sudden traffic spikes, and the operations team observes that new instances take too long to start, causing elevated latency and some request failures. The service has a large container image and initializes database connection pools at startup. Which configuration change should the team make to reduce cold-start impact while controlling cost?

A.Enable Cloud CDN for the Cloud Run service and set a long cache TTL for order-processing responses.
B.Move the service to a GKE cluster with cluster autoscaling and a horizontal pod autoscaler.
C.Set the minimum number of instances to a value greater than zero and enable CPU always allocated for the service.
D.Increase the maximum number of instances and set the container concurrency to one.
AnswerC

Setting a minimum instance count keeps warm instances ready to serve traffic, eliminating cold starts for the baseline load. Enabling CPU always allocated ensures those instances retain CPU outside request processing, which is necessary for background initialization and connection pool maintenance. Together they reduce latency during spikes while allowing the maximum instance count to scale for peak demand.

Why this answer

Cold starts occur when Cloud Run must start a new instance, and the large image plus startup initialization makes this slow. Keeping a minimum number of instances warm removes startup latency for baseline traffic, and allocating CPU outside requests lets those instances maintain connection pools. The service can still scale to the maximum instance count during peaks, so cost stays proportional to actual demand beyond the warm baseline.

Exam trap

The trap here is trying to solve startup latency by increasing maximum instances or concurrency settings, which affect scaling capacity rather than the time a new instance needs to become ready.

67
MCQeasy

Your organization requires that all production changes to Google Cloud resources be auditable and that you can identify who made a change and when. You need to configure logging to meet this requirement. What should you do?

A.Enable Admin Activity audit logs, which are enabled by default, and export them to a centralized logging project or Cloud Storage bucket with retention policies.
B.Configure VPC Flow Logs to capture network traffic and analyze it for unauthorized changes.
C.Enable Data Access audit logs for all services and export them to Cloud Storage for long-term retention.
D.Use Cloud Monitoring to create alerting policies for resource changes and send notifications to a team email.
AnswerA

Admin Activity audit logs record administrative changes to resources and are enabled by default. They include information about who made the change, what was changed, and when. Exporting them to a centralized location ensures long-term retention and auditability. This meets the requirement to identify who made a change and when.

Why this answer

Admin Activity audit logs are enabled by default and record administrative changes, including the identity of the caller and the timestamp. Exporting these logs to a centralized project or Cloud Storage with retention policies ensures they are preserved for auditing. This satisfies the requirement to audit production changes.

Exam trap

The trap here is confusing Data Access audit logs with Admin Activity audit logs; the former are for data reads/writes and are not enabled by default.

68
MCQmedium

Your company uses Cloud VPN (HA VPN) to connect to Google Cloud. You need to achieve a 99.99% SLA for the VPN connection. What configuration is required?

A.One VPN gateway with four tunnels to different on-premises devices
B.Two VPN gateways, each with two tunnels, totaling four tunnels
C.Two VPN gateways, each with one tunnel, using two different edge availability domains
D.One VPN gateway with two tunnels to the same on-premises device
AnswerB

Four tunnels across two HA VPN gateways satisfy the 99.99% SLA, since Google requires at least two tunnels on distinct gateways to guarantee that tier. Each gateway provides redundancy, so a single gateway or tunnel failure does not drop the connection, meeting the stem's availability constraint.

Why this answer

To achieve the 99.99% SLA for HA VPN, Google Cloud requires two VPN gateways, each with two tunnels, for a total of four tunnels. This configuration provides redundancy across both gateways and tunnels, satisfying the availability requirement defined by Google's HA VPN SLA. A single gateway, even with multiple tunnels, cannot meet the 99.99% SLA.

Exam trap

The trap is assuming that more tunnels on a single gateway equals higher availability — candidates must remember that the 99.99% SLA specifically requires two gateways with two tunnels each, not just four tunnels anywhere.

How to eliminate wrong answers

Option A is wrong because a single VPN gateway with four tunnels does not provide gateway-level redundancy — if the gateway fails, all tunnels fail, capping the SLA at 99.9%. Option C is wrong because two gateways with only one tunnel each provides gateway redundancy but not tunnel redundancy within each gateway; the 99.99% SLA requires two tunnels per gateway. Option D is wrong because a single gateway with two tunnels to the same on-premises device offers no gateway redundancy and no peer redundancy, yielding at most 99.9%.

69
MCQhard

Your organization runs a critical application on Google Cloud that uses Cloud SQL for PostgreSQL. The database is in us-central1. The business requires a recovery point objective (RPO) of 5 minutes and a recovery time objective (RTO) of 1 hour in case of a regional failure. What should you do?

A.Use Cloud SQL point-in-time recovery (PITR) to restore to a specific time in another region.
B.Configure a cross-region read replica and promote it in case of regional failure.
C.Enable high availability (HA) for the Cloud SQL instance, which provides a standby in another zone.
D.Export the database to Cloud Storage every 5 minutes and import it into a new instance in another region during a disaster.
AnswerB

A cross-region read replica replicates asynchronously to another region. In a regional failure, you can promote the replica to a standalone instance. This provides an RPO of typically less than 5 minutes and an RTO of under 1 hour, meeting the requirements. It is the recommended approach for regional disaster recovery.

Why this answer

A cross-region read replica asynchronously replicates data to a different region. In a regional failure, promoting the replica provides a recovery point close to the failure time (RPO within minutes) and can be done quickly (RTO under 1 hour). This is the standard solution for regional DR with Cloud SQL.

Exam trap

The trap here is assuming that HA protects against regional failures, but HA only provides zonal redundancy within a region.

70
MCQmedium

An organization wants to receive alerts when their Cloud SQL instance's CPU utilization exceeds 80% for 5 minutes. They want to send the alert to both email and a Pub/Sub topic for further processing. What should they do?

A.Configure a Cloud Scheduler job to check CPU utilization and publish to Pub/Sub
B.Create a log-based alert for CPU utilization using Logging and route to email and Pub/Sub
C.Create a Cloud Monitoring alerting policy with a metric threshold condition on CPU utilization and add both email and Pub/Sub notification channels
D.Use Cloud Functions to poll the Cloud Monitoring API every minute and send notifications
AnswerC

A Cloud Monitoring alerting policy with a metric threshold condition evaluates CPU utilisation against 80% for the five-minute duration. Attaching both email and Pub/Sub notification channels delivers the alert to each destination, satisfying the dual-delivery requirement without custom code.

Why this answer

Cloud Monitoring alerting policies support metric threshold conditions (e.g., CPU utilization > 80% for 5 minutes) and allow multiple notification channels, including email and Pub/Sub. Creating an alerting policy with a metric threshold condition on the Cloud SQL CPU metric and adding both email and Pub/Sub channels satisfies the requirement directly. This is the native, event-driven approach in Google Cloud.

Exam trap

PCA often tests whether candidates confuse log-based alerts (which trigger on log entries) with metric-based alerting policies (which trigger on metric thresholds), leading them to pick the log-based option for a CPU metric.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler is a cron service, not a monitoring/alerting system, and polling CPU utilization manually is inefficient and not the intended design. Option B is wrong because log-based alerts trigger on log entries, not on metric thresholds like CPU utilization, so they cannot directly alert on a CPU metric. Option D is wrong because polling the Monitoring API with Cloud Functions is a custom, fragile workaround that duplicates functionality already provided by alerting policies.

71
MCQmedium

A company has a Cloud SQL for MySQL instance with automated backups enabled. They need to recover the database to a specific point in time within the last hour. Which feature should they use?

A.Failover replica
B.Point-in-time recovery (PITR)
C.Automated backup restore
D.Import using the mysqldump file
AnswerB

Point-in-time recovery uses binary logs to restore a Cloud SQL for MySQL instance to a specific timestamp, not just the last automated backup. This satisfies the requirement to recover to a point within the last hour.

Why this answer

Point-in-time recovery (PITR) lets Cloud SQL for MySQL restore to a specific timestamp within the retention window by combining automated backups with binary logs. It is the only option that supports recovery to an arbitrary point within the last hour. Automated backup restore only returns the database to the time of the last backup, not an arbitrary point.

Exam trap

The trap here is confusing high-availability features like failover replicas with backup and recovery features; PCA candidates often pick failover replica when the scenario is about restoring to a past point in time.

How to eliminate wrong answers

Option A is wrong because a failover replica is for high availability during a zone or instance failure, not for recovering to a past point in time. Option C is wrong because restoring an automated backup recovers only to the backup's creation time, which may be hours old and cannot target a specific minute. Option D is wrong because mysqldump is a logical export/import tool and is not the mechanism for point-in-time recovery; it also requires a pre-existing dump file.

72
MCQmedium

An organization uses Cloud Storage to store critical documents. They want to protect against accidental deletion or overwriting of objects. Which feature should they enable?

A.Uniform bucket-level access
B.Object lifecycle management rules
C.Object versioning and retention policies
D.Customer-managed encryption keys (CMEK)
AnswerC

Versioning preserves every prior generation of an object, so an overwrite creates a new version rather than destroying the original, and deletion only adds a delete marker. Retention policies add immutability, satisfying the requirement to protect critical documents from accidental deletion or overwriting.

Why this answer

Object versioning and retention policies together protect against accidental deletion and overwrites. Versioning keeps multiple versions of objects, and retention policies prevent deletion until a specified time. Uniform bucket-level access is for access control, not protection.

Object lifecycle management automates transitions/deletion, not protection. Encryption protects data at rest.

73
MCQmedium

Your company has a production Cloud SQL for PostgreSQL instance in us-central1 with automated backups enabled. You need to ensure that if the zone fails, the database automatically fails over to a standby in a different zone with minimal downtime. What should you do?

A.Enable deletion protection on the instance.
B.Create a cross-region read replica and manually promote it during a failure.
C.Configure the instance as a highly available (regional) instance.
D.Enable point-in-time recovery (PITR) and keep 30 days of transaction logs.
AnswerC

Configuring a highly available (regional) instance provisions a standby in a different zone within the same region, with automatic failover and synchronous replication. This satisfies the stem's requirement for automatic zone-failure failover with minimal downtime, which a single-zone instance with backups alone cannot provide.

Why this answer

Configuring the Cloud SQL instance as a highly available (regional) instance provisions a standby in a different zone within the same region and automatically fails over during a zone failure with minimal downtime. This is the native HA mechanism for Cloud SQL and directly satisfies the requirement.

Exam trap

PCA often tests the difference between HA (automatic zone failover) and read replicas (manual promotion for DR) — candidates pick the cross-region replica thinking it provides automatic failover, but it requires manual intervention.

How to eliminate wrong answers

Option A is wrong because deletion protection only prevents accidental instance deletion — it has nothing to do with zone failover or availability. Option B is wrong because a cross-region read replica requires manual promotion and is designed for regional disaster recovery, not automatic zone-level failover with minimal downtime. Option D is wrong because point-in-time recovery restores data to a prior moment by replaying transaction logs — it is a data recovery feature, not an availability or failover mechanism.

74
MCQmedium

Your team operates a production e-commerce application on a managed instance group (MIG) that serves traffic through a global external Application Load Balancer. During a new release, the team wants to deploy the new version to a small subset of instances and then progressively increase traffic to it while monitoring error rates, with the ability to immediately roll back if errors spike. The new version is already built as a custom image. Which approach should you use?

A.Create a new global external Application Load Balancer with a separate backend service pointing only to the new image, and use Cloud DNS weighted routing to send 10% of users to the new load balancer.
B.Create a second MIG with the new image, add it as a backend to the existing backend service with a small capacity, and gradually shift traffic between the two MIGs using weighted traffic distribution in the backend service.
C.Perform an in-place update of the MIG template to the new image with a very small maxUnavailable, and rely on the load balancer health checks to remove unhealthy instances automatically.
D.Use a rolling update with maxSurge and maxUnavailable set to 50% so that half the instances are replaced at once, then wait for health checks to pass before continuing.
AnswerB

Weighted traffic distribution on a backend service lets you send a controlled percentage of user traffic to a new MIG while keeping the rest on the stable version. Monitoring error rates and adjusting weights gives progressive rollout and instant rollback by setting the new backend weight to zero. This matches canary release requirements without rebuilding instances.

Why this answer

The requirement is a controlled canary with progressive traffic shifting and fast rollback. Weighted traffic distribution on a backend service allows two MIGs, each running a different image, to receive defined percentages of live traffic. You can start small, watch error rates, increase the weight, and set it back to zero if problems appear.

Other approaches either replace instances in place or rely on DNS, which lacks the precision and quick reversibility needed.

Exam trap

The trap here is assuming that a rolling update with maxSurge and maxUnavailable is equivalent to a canary release, when rolling updates replace instances in place and cannot route a precise percentage of user traffic to a new version.

75
Multi-Selecthard

Your company wants to implement a canary deployment for a microservice running on GKE. You need to gradually shift traffic from the stable version to the canary version while monitoring error rates. Which THREE components or practices should you use? (Choose 3)

Select 3 answers
A.Cloud Deploy with an automated canary strategy and verification
B.Cloud Monitoring to track error rates and trigger rollback
C.Cloud CDN for caching responses
D.Feature flags in the application code
E.Istio for traffic splitting between versions
AnswersA, B, E

Cloud Deploy's automated canary strategy progressively shifts traffic percentages between GKE revisions and runs verification steps, halting or rolling back when analysis fails. It directly provides the staged rollout and metric-gated promotion the scenario demands.

Why this answer

Option A is correct because Cloud Deploy natively supports canary deployment strategies with configurable phases (for example, 50% then 100%) and automated verification that can advance or halt a rollout based on analysis results. Option B is correct because Cloud Monitoring collects the error-rate metrics (such as HTTP 5xx ratios from the service) that Cloud Deploy's verification step or alerting policies use to detect failures and trigger a rollback. Option E is correct because Istio on GKE provides fine-grained traffic splitting via VirtualService weights, letting you shift a precise percentage of requests from the stable to the canary version while observing behavior.

Option C is not appropriate because Cloud CDN caches responses at the edge and does not perform version-based traffic shifting or canary analysis. Option D is not appropriate because feature flags toggle functionality inside a single deployed version and do not by themselves implement gradual traffic shifting between two separately deployed versions.

Exam trap

PCA often tests the confusion between feature flags and canary deployments; candidates pick feature flags because they sound like gradual rollout, but feature flags do not shift traffic between deployed versions.

Page 1 of 2 · 78 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Managing Implementation and Ensuring Solution and Operations Reliability questions.