Courseiva

CCNA Ace Ensuring Operation Questions

75 of 81 questions · Page 1/2 · Ace Ensuring Operation topic · Answers revealed

1
Multi-Selectmedium

You want to create a log-based metric to count errors from your application logs. Which TWO resources are required? (Select 2)

Select 2 answers
A.A filter that matches the error log entries
B.An alerting policy
C.A metric descriptor (e.g., name, type, label)
D.A log sink
E.A notification channel
AnswersA, C

In Cloud Logging, a logs-based metric is created by defining a filter that selects which log entries increment the metric's counter. This filter is the heart of the metric because it evaluates each incoming log entry against conditions such as severity >= ERROR or a text payload match. Without it, the metric has no way to distinguish error logs from other entries, so the filter directly determines the metric's value and is therefore required.

Why this answer

You need a filter to match error logs and a metric descriptor that defines the metric type.

2
MCQmedium

You have updated a deployment in GKE, but the new pods are crashing. You want to revert to the previous working version. What should you do?

A.kubectl rollout status deployment/my-app
B.kubectl rollout undo deployment/my-app
C.kubectl scale deployment/my-app --replicas=0
D.kubectl delete deployment/my-app and recreate
AnswerB

This command reverts the Deployment to the previous revision by rolling back to the last good ReplicaSet. The Deployment controller will scale down the current ReplicaSet and scale up the old one, restoring the previous container image and configuration. This is the correct, built-in way to undo a bad deployment while maintaining availability.

Why this answer

kubectl rollout undo reverts to the previous revision.

3
MCQeasy

You are using Cloud Run and want to split traffic so that 10% of requests go to revision v2 and 90% go to revision v1. Which command should you use?

A.gcloud run deploy --image my-image --traffic v1=90,v2=10
B.gcloud run services update --traffic v1=90,v2=10
C.gcloud run revisions update v2 --traffic 10
D.gcloud run services update-traffic --to-revisions v1=90,v2=10
AnswerD

This is the correct command for splitting traffic between already deployed revisions: `gcloud run services update-traffic` with `--to-revisions` takes a comma-separated list of `revision=percentage` pairs (v1=90,v2=10) and applies the routing immediately. The specified revisions must exist and the percentages must total 100. It does not create a new revision, so it is the appropriate operation after v1 and v2 have both been deployed.

Why this answer

The correct command to split traffic between Cloud Run revisions is `gcloud run services update-traffic` with the `--to-revisions` flag, specifying the revision names and percentages. This command updates the traffic allocation for a service. The syntax `--to-revisions v1=90,v2=10` correctly assigns 90% to v1 and 10% to v2.

Exam trap

ACE often tests the exact gcloud command syntax for traffic splitting; candidates may confuse `gcloud run deploy` with `gcloud run services update-traffic`, or use incorrect flags like `--traffic` instead of `--to-revisions`.

How to eliminate wrong answers

Option A is wrong because `gcloud run deploy` is used to deploy a new revision, and the `--traffic` flag is not valid for splitting traffic in that manner; it is used for setting traffic during deployment but not with revision names like that. Option B is wrong because `gcloud run services update` does not have a `--traffic` flag; traffic splitting is done via `update-traffic`. Option C is wrong because `gcloud run revisions update` is not a valid command for traffic management; revisions are immutable and traffic is managed at the service level.

4
Multi-Selectmedium

An engineer needs to create a Cloud Monitoring dashboard that displays CPU utilization for all Compute Engine instances in a project. Which TWO steps are required? (Choose 2)

Select 2 answers
A.Create an uptime check
B.Add the chart to a dashboard
C.Create a chart using Metric Explorer
D.Create a log-based metric
E.Set up a notification channel
AnswersB, C

After you generate a chart in Metric Explorer, adding it to a dashboard persists the visualization as a widget in a chosen layout, making it visible to the team and available in the Monitoring UI. This is the final, required step to actually place the metric on the dashboard; without it, the chart exists only in the temporary Metric Explorer session and will be lost when you navigate away.

Why this answer

First, use Metric Explorer to create a chart with the CPU utilization metric. Then, add that chart to a dashboard. Dashboards can have charts from Metric Explorer.

You do not need to create an alert or export logs.

5
MCQmedium

You have a Cloud Storage bucket that contains sensitive data. You need to ensure that all access to the bucket is logged, including data reads and writes, and that the logs are retained for at least one year. You also want to be able to analyze the logs using BigQuery. What should you do?

A.Use Cloud Audit Logs to capture admin activity only, and set up a log sink to BigQuery with a retention period of 365 days.
B.Create a log sink to Cloud Storage, and then use a BigQuery external table to query the logs directly from Cloud Storage.
C.Enable Cloud Storage bucket logging by using the gsutil logging set command, and then export the logs to BigQuery using a log sink.
D.Enable Data Access audit logs for Cloud Storage, create a log sink to BigQuery, and set the retention period on the BigQuery dataset to 365 days.
AnswerD

Data Access audit logs capture read and write operations on Cloud Storage. Enabling them ensures all access is logged. A log sink to BigQuery exports the logs for analysis, and setting the dataset's default table expiration to 365 days ensures retention for at least one year. This meets all requirements.

Why this answer

To log all access to a Cloud Storage bucket, including data reads and writes, you must enable Data Access audit logs. Then, create a log sink to BigQuery for analysis and set the dataset's retention to 365 days. Legacy bucket logging or Admin Activity logs alone do not capture all data access events.

Exam trap

The trap here is assuming that Admin Activity audit logs include data reads and writes, when in fact Data Access audit logs are needed for that level of detail.

6
MCQhard

You have a Cloud Run service that experiences intermittent high latency. You want to analyze the latency of specific request paths to identify bottlenecks. You enable Cloud Trace and instrument your application with OpenTelemetry. Which tool or feature should you use to view a waterfall diagram of latencies across services for a single request?

A.Error Reporting
B.Cloud Trace Trace List and Trace Details
C.Cloud Monitoring Metrics Explorer
D.Cloud Logging Logs Explorer
AnswerB

Cloud Trace Trace List and Trace Details is the correct service because the Trace List displays each sampled request as a row with its overall latency, while Trace Details opens a waterfall chart that breaks the request into individual spans. In a Cloud Run service, this shows time spent in container startup, internal logic, and downstream calls, making it possible to pinpoint exactly which span causes an intermittent slowdown. The per-request, span-level granularity directly matches the need to diagnose variable performance.

Why this answer

Cloud Trace Trace List and Trace Details provide a waterfall diagram that visualizes the latency of each request as it propagates through different services. This allows you to pinpoint bottlenecks by seeing the duration of each span in the trace. OpenTelemetry instrumentation enriches traces with custom spans, making this the correct tool for analyzing request paths.

Exam trap

The trap is confusing Cloud Trace with Cloud Monitoring or Logging; candidates might think Metrics Explorer or Logs Explorer can show waterfall diagrams, but only Cloud Trace provides that view.

How to eliminate wrong answers

Option A is wrong because Error Reporting aggregates and displays errors, not latency waterfalls. Option C is wrong because Metrics Explorer is for viewing and alerting on metrics, not for tracing individual requests. Option D is wrong because Logs Explorer is for searching and analyzing logs, not for visualizing trace waterfalls.

7
MCQeasy

You have a Compute Engine instance that is running a CPU-intensive workload. After monitoring, you realize the machine type needs to be upgraded to a larger CPU. What is the correct sequence to change the machine type?

A.Stop the instance, run gcloud compute instances set-machine-type, then start the instance
B.Run gcloud compute instances set-machine-type while the instance is running
C.Delete the instance and create a new one with the desired machine type
D.Use gcloud compute instances update to change the machine type
AnswerA

The correct sequence is to first stop the instance with `gcloud compute instances stop INSTANCE_NAME`, then run `gcloud compute instances set-machine-type INSTANCE_NAME --machine-type MACHINE_TYPE` while the instance is in the TERMINATED state, and finally start it again with `gcloud compute instances start INSTANCE_NAME`. This preserves the instance's boot disk, persistent disks, static IP, metadata, and other configuration, and is the standard non-destructive way to resize a VM.

Why this answer

Changing a Compute Engine instance's machine type requires the instance to be stopped first, because the machine type determines the underlying host resources. The correct sequence is to stop the instance, run `gcloud compute instances set-machine-type`, then start the instance.

Exam trap

The trap is assuming machine type can be changed live like a CPU hot-add — candidates pick the running-instance option, forgetting that GCE requires the instance to be stopped (TERMINATED) before the machine type can be modified.

How to eliminate wrong answers

Option B is wrong because `set-machine-type` cannot be applied to a running instance — the API returns an error requiring the instance to be TERMINATED. Option C is wrong because deleting and recreating the instance is unnecessary and risks data loss unless disks are preserved; the in-place method is preferred. Option D is wrong because `gcloud compute instances update` modifies properties like labels or metadata, not the machine type.

8
Multi-Selecteasy

You want to monitor the uptime of an external HTTP endpoint from multiple locations around the world. Which TWO steps should you take? (Choose two.)

Select 2 answers
A.Create a log-based metric for the endpoint response time
B.Create a notification channel of type Pub/Sub
C.Select multiple locations (e.g., us-west1, europe-west1, asia-east1) for the uptime check
D.Create an uptime check in Cloud Monitoring with the HTTP endpoint URL
E.Enable VPC flow logs for the endpoint
AnswersC, D

Choosing several geographic locations makes the uptime check probe the endpoint from distinct regions, so an outage or regional network fault is distinguished from a localised failure. This directly satisfies the requirement to monitor availability from multiple locations around the world.

Why this answer

Option D is correct because a Cloud Monitoring uptime check is the native resource designed to probe an HTTP(S) endpoint from Google-managed probers and report up/down status and latency. Option C is correct because an uptime check lets you select multiple geographic regions (e.g., us-west1, europe-west1, asia-east1) so the endpoint is tested from several locations worldwide, which is exactly the requirement. Option A is not appropriate because a log-based metric derives values from log entries and does not itself perform external HTTP probing from multiple regions.

Option B is not appropriate because a Pub/Sub notification channel only delivers alerts; it does not create the monitoring check or provide multi-region probing. Option E is not appropriate because VPC flow logs capture IP traffic metadata within a VPC and cannot monitor an external HTTP endpoint's uptime.

Exam trap

ACE often tests the trap of confusing uptime checks with logging or notification mechanisms, leading candidates to select log-based metrics or Pub/Sub channels instead of the core uptime check configuration.

9
MCQeasy

A developer needs to query BigQuery using the bq command-line tool with standard SQL. Which flag should they include?

A.--format
B.--use_legacy_sql=false
C.--project_id
D.--sync
AnswerB

`--use_legacy_sql=false` is the required flag because the `bq` command-line tool historically defaults to legacy SQL when running queries. Legacy SQL uses a different syntax and operates differently from standard GoogleSQL (e.g., unique handling of JOINs and functions). Passing `false` explicitly switches the parser to standard SQL, allowing the developer's query to run without rewriting it into legacy dialect.

Why this answer

The '--use_legacy_sql=false' flag enables standard SQL. By default, bq uses legacy SQL. '--format' controls output format, not SQL dialect. '--project_id' specifies project. '--sync' is not a valid bq flag.

10
MCQhard

You need to collect and analyze latency traces for a microservices application running on GKE. You want to identify which services are contributing to overall latency. Which Google Cloud service should you enable and use?

A.Cloud Logging
B.Cloud Profiler
C.Cloud Monitoring
D.Cloud Trace
AnswerD

Cloud Trace is the Google Cloud service specifically designed for distributed tracing: it captures spans from instrumented applications or via OpenTelemetry, groups them into traces for each request, and renders a waterfall view showing where time is spent across microservices. It supports latency distribution analysis, allows comparison of recent traces, and can identify bottleneck services and anomalously slow requests. Thus it directly answers the need to collect and analyze latency traces.

Why this answer

Cloud Trace is Google Cloud's distributed tracing service, designed to collect latency data across microservices and visualize trace waterfalls so you can pinpoint which service contributes to end-to-end latency. It integrates natively with GKE and supports OpenTelemetry and the Cloud Trace agent.

Exam trap

The trap is conflating Monitoring (metrics/alerts) with Trace (distributed tracing) — candidates pick Cloud Monitoring because it 'shows latency graphs', but only Trace provides per-service span attribution.

How to eliminate wrong answers

Option A is wrong because Cloud Logging captures log events, not latency traces or span-level timing. Option B is wrong because Cloud Profiler analyzes CPU and memory usage of running code, not request latency across services. Option C is wrong because Cloud Monitoring collects metrics and can alert on latency, but it does not provide distributed trace waterfalls for per-service latency attribution.

11
MCQhard

You are deploying a GKE cluster with node autoscaling enabled. The cluster runs batch jobs that are sensitive to startup latency. You notice that during scale-up, new nodes take several minutes to become ready. Which action can reduce the time it takes for new nodes to join the cluster?

A.Increase the initial node pool size
B.Set the --max-nodes-per-pool flag to a higher value
C.Use a custom image with pre-installed dependencies
D.Enable cluster autoscaler with --enable-autorepair
AnswerC

Using a custom image with pre-installed dependencies is the correct approach because it directly reduces node initialization time. A custom image can bake in the container runtime, required OS packages, and even pre-cached application container images, avoiding the typical runtime download and configuration steps when a new node is added. When the cluster autoscaler triggers a scale-out, these nodes become schedulable faster, so pending pods are scheduled more quickly.

Why this answer

Using a custom image with pre-installed dependencies reduces the time new nodes spend pulling and installing container images or configuring software after the node boots. By baking dependencies into the node image, the node can join the cluster and start running workloads faster, directly addressing the startup latency issue for batch jobs.

Exam trap

ACE often tests the misconception that increasing node pool size or max-nodes limits reduces startup latency, when the actual fix is optimizing the node image and pre-installing dependencies.

How to eliminate wrong answers

Option A is wrong because increasing the initial node pool size only provides more capacity upfront; it does not reduce the time it takes for individual new nodes to become ready during scale-up. Option B is wrong because --max-nodes-per-pool controls the upper limit of nodes in a pool, not the startup time of individual nodes; raising it does not speed up node readiness. Option D is wrong because --enable-autorepair enables automatic repair of unhealthy nodes, which is unrelated to reducing node startup latency and may even add overhead.

12
MCQmedium

Your BigQuery query is taking longer than expected. You want to estimate the query cost before running it and get a preview of how many bytes will be processed. Which bq command should you use?

A.bq show --format=prettyjson mydataset.mytable
B.bq ls --format=prettyjson mydataset
C.bq query --use_legacy_sql=false --dry_run 'SELECT ...'
D.bq query --use_legacy_sql=false --batch 'SELECT ...'
AnswerC

bq query --use_legacy_sql=false --dry_run 'SELECT ...' sends the query to BigQuery's planner, which validates the SQL and returns the estimated number of bytes that would be read from storage, without actually executing the query or consuming slots. The --use_legacy_sql=false flag ensures your statement is parsed as standard SQL, not the older legacy dialect, which matters for syntax compatibility. This is exactly the right tool when a query is slow and you want to quickly see how much data it touches before investing time in optimization or running it.

Why this answer

The bq query --dry_run flag validates the query and returns the estimated number of bytes that will be processed without actually executing it or incurring charges. This is the standard way to preview BigQuery query cost before running it, since on-demand pricing is based on bytes scanned. The --use_legacy_sql=false flag ensures standard SQL syntax is used.

Exam trap

The trap is assuming any bq command that inspects a table or query gives cost information — candidates often pick bq show because it returns metadata, not realizing only --dry_run provides the bytes-processed estimate without executing the query.

How to eliminate wrong answers

Option A is wrong because bq show displays table metadata (schema, partitioning, row count) but does not estimate query cost or bytes processed. Option B is wrong because bq ls lists datasets or tables in a dataset; it provides no query cost information. Option D is wrong because --batch runs the query asynchronously in the background but still executes it and incurs charges — it does not provide a cost preview.

13
MCQmedium

A company wants to split traffic between two revisions of a Cloud Run service: 90% to revision 'green' and 10% to revision 'blue'. Which command should they use?

A.gcloud run revisions list
B.gcloud run services update
C.gcloud run services update-traffic
D.gcloud run deploy
AnswerC

`gcloud run services update-traffic` is the correct command to split traffic between two or more existing revisions of a Cloud Run service. It accepts flags like `--to-revisions=rev1=50,rev2=50` to assign precise percentages, or `--to-latest` to route all traffic to the latest revision. This command directly modifies the route resource, making it the appropriate tool for controlled canary rollouts or rollbacks.

Why this answer

'gcloud run services update-traffic' is the correct command to manage traffic splitting between revisions. 'gcloud run revisions list' only lists revisions. 'gcloud run services update' does not handle traffic directly. 'gcloud run deploy' with --no-traffic is for initial deployment.

14
MCQhard

An engineer needs to update a Kubernetes Deployment's container image to version v2. They run 'kubectl set image deployment/my-app my-container=gcr.io/my-project/my-image:v2'. After a few minutes, they check the rollout status and see a failure. They want to revert to the previous image. Which command should they use?

A.kubectl rollout status deployment/my-app
B.kubectl rollout undo deployment/my-app
C.kubectl delete deployment/my-app --cascade=false
D.kubectl set image deployment/my-app my-container=gcr.io/my-project/my-image:v1
AnswerB

kubectl rollout undo deployment/my-app is the correct command because it reverts the deployment to the previous revision, restoring the prior pod template spec and container image. Kubernetes retains rollout history for each change to the pod template, and undo automatically scales down the current ReplicaSet and scales up the previous one, seamlessly rolling back the application without manual image specification.

Why this answer

'kubectl rollout undo' reverts the Deployment to the previous revision. 'kubectl rollout status' shows status but does not revert. 'kubectl set image' with v1 would manually set the old image, but 'undo' is the standard rollback command.

15
MCQmedium

You have a Cloud Run service that is experiencing high latency. You want to analyze the latency distribution of requests. Which Google Cloud tool should you use?

A.Cloud Debugger
B.Cloud Logging Log Explorer
C.Cloud Trace
D.Cloud Monitoring Metrics Explorer
AnswerC

Cloud Trace is purpose-built for latency analysis. It collects latency data from Cloud Run and other GCP services, then generates distributed traces with spans that show the duration of each operation—such as receiving the request, calling downstream dependencies, and returning the response. Trace features like waterfall views, latency distributions, and per-trace breakdowns let you identify exactly which service or API call is the bottleneck, making it the correct tool for high-latency issues.

Why this answer

Cloud Trace is a distributed tracing service that collects latency data from applications and provides detailed analysis, including latency distributions and per-request traces.

16
MCQeasy

You have a Pub/Sub subscription that is accumulating a backlog of messages. Which Cloud Monitoring metric should you alert on to detect this condition?

A.pubsub.googleapis.com/subscription/oldest_unacked_message_age
B.pubsub.googleapis.com/subscription/sent_messages_count
C.pubsub.googleapis.com/subscription/unacked_messages_by_region
D.pubsub.googleapis.com/subscription/ack_message_count
AnswerA

This metric tracks the maximum age of the oldest message that has not yet been acknowledged by any subscriber for the subscription. It directly reflects backlog depth and consumer lag: when a subscription is accumulating a backlog, this value grows steadily because messages sit unacked for longer periods. It is the ideal signal for alerting on message processing delays because it captures the time dimension of the backlog, not just its size.

Why this answer

The metric oldest_unacked_message_age measures the age of the oldest unacknowledged message in a subscription. A growing backlog directly causes this age to increase, making it the most reliable indicator of a subscription falling behind. Alerting on this metric allows you to detect when messages are not being processed in a timely manner.

Exam trap

ACE often tests the difference between rate metrics (sent, ack) and backlog age metrics; candidates may mistakenly choose ack_message_count thinking it reflects backlog, but it does not.

How to eliminate wrong answers

Option B is wrong because sent_messages_count only shows the rate of messages sent to the subscription, not whether they are being acknowledged. Option C is wrong because unacked_messages_by_region is not a standard Cloud Monitoring metric for Pub/Sub; the correct metric is unacked_messages, but it does not directly indicate backlog age. Option D is wrong because ack_message_count measures the rate of acknowledgments, which could be high even if a backlog exists due to a sudden spike in incoming messages.

17
Multi-Selecthard

Your application running on Compute Engine is experiencing intermittent high latency. You need to diagnose the root cause. Which THREE tools or services should you use to gather data? (Choose 3)

Select 3 answers
A.Cloud Monitoring
B.Cloud Logging
C.Cloud Profiler
D.Cloud Debugger
E.Cloud Trace
AnswersA, B, E

Cloud Monitoring is the correct starting point because it provides time-series metrics for Compute Engine, such as CPU utilization, memory usage, disk I/O, and network throughput. You can build custom dashboards and alerts to correlate intermittent latency spikes with resource saturation, helping you determine whether the cause is a bottleneck in the VM, disk, or network. This metric-centric view is essential for seeing the pattern of when slowdowns occur.

Why this answer

Cloud Monitoring provides metrics and dashboards; Cloud Logging provides logs; Cloud Trace provides trace data for latency analysis. Together they cover metrics, logs, and traces for comprehensive troubleshooting.

18
MCQmedium

A company wants to export all Cloud Logging logs to BigQuery for long-term analysis. They create a log sink with a BigQuery dataset as the destination. After a few days, they notice that some logs are missing in BigQuery. What is the most likely reason?

A.The sink's inclusion filter is too restrictive
B.Logs older than 30 days cannot be exported
C.The sink's destination is a table, not a dataset
D.BigQuery dataset is in a different region
AnswerA

This is the correct diagnostic. A log sink only forwards entries that match its inclusion filter, and an overly narrow filter—such as one restricted to a single resource type or severity level—will silently exclude the rest of the log stream before it reaches BigQuery. Check the sink's filter in the Logs Explorer to confirm it matches the actual log entries you expect to export, and note that any exclusion filters are applied after the inclusion filter and can further reduce the data routed.

Why this answer

Log sinks have a buffer period of up to a few minutes, but they guarantee delivery. However, if the sink's filter excludes certain logs (e.g., by resource type or severity), those logs are not exported. Missing logs usually indicate a filter misconfiguration.

19
MCQhard

Your GKE cluster nodes are running low on resources. You need to enable node pool autoscaling so that the cluster automatically adds and removes nodes based on demand. The node pool is named 'default-pool'. Which command completes this task?

A.gcloud container node-pools update default-pool --autoscaling enabled
B.gcloud container node-pools update default-pool --enable-autoscaling --min-nodes 1 --max-nodes 10
C.kubectl autoscale node-pool default-pool --min 1 --max 10
D.gcloud container clusters update my-cluster --enable-autoscaling --min-nodes 1 --max-nodes 10
AnswerB

This is the correct gcloud command to enable cluster autoscaler on a specific node pool. The `--enable-autoscaling` switch turns on autoscaling for the `default-pool`, and the `--min-nodes 1` and `--max-nodes 10` flags define the minimum and maximum size of the node pool. With this configuration, GKE's cluster autoscaler will automatically add or remove nodes within that range based on pending pod resource requests, which directly relieves the low-resource condition on the cluster's nodes.

Why this answer

gcloud container node-pools update with --enable-autoscaling enables autoscaling, and --min-nodes/--max-nodes set boundaries.

20
MCQeasy

An engineer needs to create an alerting policy in Cloud Monitoring that sends a notification when the 99th percentile latency of a service exceeds 500 ms for 5 minutes. Which metric type should they use?

A.Metric threshold
B.Log-based metric
C.Uptime check
D.Cloud Audit Logs
AnswerA

Metric threshold is the standard condition type used in a Cloud Monitoring alerting policy. It evaluates a metric stream (e.g., Compute Engine CPU utilization, disk bytes used) against a numeric threshold over a specified aggregation window, triggering notifications when the value crosses the threshold. This is the correct answer because it directly defines the alerting condition that the engineer needs.

Why this answer

A metric threshold alert uses a numeric metric and triggers when the value crosses a threshold. Log-based alerts are for when a specific log entry appears. Uptime checks monitor availability, not latency percentiles.

21
MCQeasy

You have a Compute Engine instance running a web server. You need to allow HTTP traffic from the internet to this instance. You have already created a firewall rule that allows ingress on tcp:80 from 0.0.0.0/0. However, the instance is not receiving any traffic. You check the instance and see it has no external IP address. What should you do to allow external traffic?

A.Assign a static external IP address to the instance.
B.Configure an HTTP(S) load balancer with the instance as a backend, and use the load balancer's IP for external traffic.
C.Add an external IP address to the instance, either ephemeral or static.
D.Create a Cloud NAT gateway and configure the instance to use it.
AnswerC

Without an external IP address, the instance cannot receive traffic from the internet. Adding an external IP (ephemeral or static) enables the instance to communicate with external clients. The firewall rule already permits ingress on tcp:80, so once the instance has an external IP, traffic will reach it. This is the correct action.

Why this answer

The instance lacks an external IP address, so it cannot receive traffic from the internet even though the firewall rule allows it. Assigning an external IP (ephemeral or static) enables inbound connectivity. Cloud NAT is for outbound only, and a load balancer, while possible, adds unnecessary complexity for this scenario.

Exam trap

The trap here is thinking that a firewall rule alone is sufficient for external access, when the instance also needs a public IP address to be reachable from the internet.

22
MCQhard

You need to change the machine type of a running Compute Engine instance from n1-standard-4 to n1-standard-8. What is the correct procedure?

A.Delete the instance and recreate it with the new machine type.
B.Run gcloud compute instances set-machine-type while the instance is running.
C.Stop the instance, run gcloud compute instances set-machine-type, then start the instance.
D.Take a snapshot, create a new instance, and attach the disk.
AnswerC

This is the only correct sequence: first stop the instance (gcloud compute instances stop), which transitions it to the TERMINATED state, then call gcloud compute instances set-machine-type with the desired type (predefined, custom, or E2), and finally start the instance again. The stop/start cycle is required because the hypervisor must release the old vCPU and memory resources before the new allocation can be applied. After the instance starts, it retains its existing disks, IP addresses, and metadata, so no configuration is lost.

Why this answer

Changing machine type requires stopping the instance first.

23
MCQmedium

You need to export all Cloud Logging logs from a specific project to BigQuery for long-term analysis. What should you create?

A.A log-based metric with BigQuery as destination
B.A log sink with BigQuery as the destination
C.A Pub/Sub subscription that pushes logs to BigQuery
D.An export job from Logging to BigQuery using gcloud logging export
AnswerB

A log sink is the correct Cloud Logging resource for streaming log entries to a supported destination, and BigQuery is a first-class destination. When you create a sink with `gcloud logging sinks create` or in the console, you specify a BigQuery dataset as the destination; Cloud Logging then continuously routes exported log entries into a partitioned table. This satisfies the requirement to export all logs from the project, optionally filtered by a log query.

Why this answer

To export logs from Cloud Logging to BigQuery, you create a log sink with BigQuery as the destination. Log sinks route log entries to supported destinations, and BigQuery is a native destination for long-term analytics.

Exam trap

ACE often tests the difference between log sinks and log-based metrics, and candidates may incorrectly choose a metric or a Pub/Sub subscription when the requirement is direct export to BigQuery.

How to eliminate wrong answers

Option A is wrong because log-based metrics are used to create metric data from logs for monitoring and alerting, not for exporting raw logs to BigQuery. Option C is wrong because a Pub/Sub subscription can push logs to other services, but it is not the direct mechanism for exporting to BigQuery; you would need an additional pipeline. Option D is wrong because there is no gcloud logging export command; the correct command is gcloud logging sinks create.

24
MCQmedium

You need to attach an existing 100 GB persistent disk named 'my-disk' to a Compute Engine instance 'web-server-1'. What is the correct command?

A.gcloud compute instances add-disk web-server-1 --disk my-disk
B.gcloud compute disks attach my-disk --instance web-server-1
C.gcloud compute disks create my-disk --instance web-server-1
D.gcloud compute instances attach-disk web-server-1 --disk my-disk
AnswerD

This is the correct command. The attach-disk subcommand belongs to gcloud compute instances, takes web-server-1 as the resource, and uses --disk my-disk to specify the existing persistent disk to attach. It associates the already provisioned 100 GB disk with the instance, and requires the disk and instance to be in the same zone (or use --zone if they are not set).

Why this answer

The command is gcloud compute instances attach-disk.

25
MCQmedium

A developer needs to check the latest logs from a Compute Engine instance to debug a failed startup script. They want to filter logs from the last hour with severity ERROR or higher. Which Cloud Logging query language filter should they use?

A.resource.type="gce_instance" timestamp>="-1h" severity>=ERROR
B.resource.type="gce_instance" severity=ERROR timestamp>="-1h"
C.resource.type="compute.googleapis.com/Instance" severity>=ERROR timestamp<"-1h"
D.resource.labels.instance_id="*" severity>=ERROR timestamp>="-3600s"
AnswerA

The filter combines the required resource type with a relative timestamp and severity comparison. `resource.type="gce_instance"` scopes results to the Compute Engine instance, `timestamp>="-1h"` restricts to the last hour, and `severity>=ERROR` returns ERROR and above, satisfying every constraint in the stem.

Why this answer

The correct Cloud Logging query language filter uses resource.type="gce_instance" to target Compute Engine instances, timestamp>="-1h" to filter the last hour, and severity>=ERROR to include ERROR and higher severities. The syntax severity>=ERROR is valid and includes ERROR, CRITICAL, ALERT, and EMERGENCY. This combination precisely meets the requirement.

Exam trap

ACE often tests the exact syntax of Cloud Logging query language, especially the difference between severity=ERROR and severity>=ERROR, and the correct resource type string for Compute Engine.

How to eliminate wrong answers

Option B is wrong because severity=ERROR only matches exactly ERROR, excluding higher severities like CRITICAL, which violates the 'or higher' requirement. Option C is wrong because resource.type="compute.googleapis.com/Instance" is not the correct monitored resource type for Compute Engine instances (it should be "gce_instance"), and timestamp<"-1h" would filter logs older than one hour, not the last hour. Option D is wrong because resource.labels.instance_id="*" is not a valid wildcard syntax in Cloud Logging query language, and timestamp>="-3600s" is not a supported relative time format (it should be "-1h" or a timestamp).

26
MCQeasy

You want to receive notifications when a specific metric exceeds a threshold. Which Cloud Monitoring resource defines the condition and the action?

A.Alerting policy
B.Dashboard
C.Uptime check
D.Notification channel
AnswerA

An alerting policy is the correct resource in Google Cloud Monitoring for triggering notifications based on a specific metric condition. It contains one or more conditions (e.g., metric crosses a threshold for a set duration) and references a notification channel to deliver the alert. Without an alerting policy, no metric evaluation or notification can occur.

Why this answer

An alerting policy defines conditions (metric threshold) and notification channels.

27
MCQmedium

You are using Cloud Logging and want to export all logs from a specific Compute Engine instance to BigQuery for long-term analysis. You create a log sink with a filter for the instance's resource type and labels. What additional step is required to complete the export?

A.Create a Cloud Pub/Sub topic and configure a push subscription
B.Create a BigQuery dataset and grant the log sink's service account the BigQuery Data Editor role
C.Create a Cloud Storage bucket as a staging location
D.Enable BigQuery's streaming buffer on the dataset
AnswerB

This is the correct approach because BigQuery must already exist as a dataset for the log sink to write into, and the sink's underlying writer identity (the service account) needs the BigQuery Data Editor role (roles/bigquery.dataEditor) on that dataset to create tables and insert log entries. You first create the dataset, then configure the log sink with BigQuery as its destination, and after the sink is created you copy its service account ID and grant that service account the required IAM role. Without that grant, the sink will fail with permission errors when trying to deliver logs to BigQuery.

Why this answer

A log sink exports logs to a destination, but for BigQuery the destination dataset must exist and the sink's service account must be granted the BigQuery Data Editor role on that dataset. Without this IAM binding, the sink cannot write to BigQuery and the export fails.

Exam trap

ACE often tests whether candidates know that log sinks require explicit IAM grants to the sink's writer identity on the destination — creating the sink alone is insufficient.

How to eliminate wrong answers

Option A is wrong because Pub/Sub is used for streaming exports to third-party or custom consumers, not for direct BigQuery exports — the sink can target BigQuery directly. Option C is wrong because a Cloud Storage bucket is only needed for Cloud Storage exports or as a staging area for certain batch workflows, not for BigQuery sinks. Option D is wrong because the streaming buffer is an internal BigQuery mechanism for handling streaming inserts — it is not a configuration step required to enable a log sink export.

28
MCQmedium

You have a Cloud Run service that you want to update to use a new container image. You also want to keep the previous revision available in case you need to roll back. Which command should you use?

A.gcloud run deploy my-service --image gcr.io/my-project/my-app:v2
B.gcloud run revisions update my-service --image gcr.io/my-project/my-app:v2
C.gcloud run services update --image gcr.io/my-project/my-app:v2
D.kubectl set image service/my-service my-app=gcr.io/my-project/my-app:v2
AnswerA

Running 'gcloud run deploy my-service --image gcr.io/my-project/my-app:v2' is the correct way to update a Cloud Run service's container image. The deploy command prepares a new immutable revision with the specified image, makes it the latest revision, and automatically routes traffic to it according to your service's traffic policy. Existing revisions are preserved in the revision history, enabling immediate rollback via 'gcloud run services update-traffic' if the new revision misbehaves.

Why this answer

The command 'gcloud run deploy my-service --image gcr.io/my-project/my-app:v2' deploys a new revision of the Cloud Run service using the specified container image. By default, Cloud Run keeps previous revisions, allowing rollback. This is the standard and correct way to update a service.

Exam trap

The trap is thinking that 'gcloud run services update' can change the image, but it cannot; only 'gcloud run deploy' can, and candidates might also confuse Cloud Run with Kubernetes commands.

How to eliminate wrong answers

Option B is wrong because 'gcloud run revisions update' does not exist; revisions are immutable. Option C is wrong because 'gcloud run services update' is used to update service configuration (e.g., memory, env vars) but not to change the container image; the correct flag for image is not supported in that command. Option D is wrong because kubectl is for Kubernetes, not Cloud Run, and Cloud Run services are not managed via kubectl.

29
MCQhard

Your application uses Pub/Sub to process orders. You notice that the subscription backlog is growing. Which tool should you use to analyze the latency of each step in the processing pipeline?

A.Cloud Monitoring
B.Cloud Profiler
C.Cloud Logging
D.Cloud Trace
AnswerD

Cloud Trace provides distributed tracing with per-span and per-service latency breakdowns, capturing the full path of a message as it flows through Pub/Sub and downstream services. Its waterfall view and latency distributions reveal exactly which step—from publish to processing—contributes the most delay. It is the only tool among these options specifically designed to analyze per-step latency in a distributed, asynchronous pipeline.

Why this answer

Cloud Trace is designed to analyze latency and performance of distributed applications, including Pub/Sub processing pipelines. It provides detailed trace data showing the time taken by each step, helping identify bottlenecks. Cloud Monitoring provides metrics but not per-step latency breakdown, while Cloud Profiler is for code-level profiling and Cloud Logging for log analysis.

Exam trap

The trap is confusing Cloud Trace with Cloud Monitoring or Cloud Logging; candidates might think Monitoring provides detailed latency breakdowns, but it only offers aggregate metrics.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring aggregates metrics and alerts but does not provide detailed per-request latency tracing. Option B is wrong because Cloud Profiler analyzes CPU and memory usage of code, not end-to-end latency. Option C is wrong because Cloud Logging captures log events but lacks built-in latency analysis and visualization.

30
Multi-Selectmedium

You are troubleshooting a slow Pub/Sub subscription. Which three steps should you take to diagnose the issue? (Choose three.)

Select 3 answers
A.Check the subscription's backlog in the Pub/Sub console or via gcloud pubsub subscriptions describe
B.Use Cloud Monitoring Metrics Explorer to view the subscription's backlog and ack messages count
C.Use Cloud Trace to analyze the latency of each Pub/Sub message
D.Use Cloud Debugger to inspect the subscriber code
E.Use Cloud Logging to check for subscriber errors or delivery failures
AnswersA, B, E

Checking the subscription's backlog, either in the Pub/Sub console or via the `gcloud pubsub subscriptions describe` command, gives a direct numeric measurement of how many messages are unacknowledged and waiting to be redelivered. A consistently growing backlog relative to the publish rate indicates the subscriber cannot keep up with the incoming flow. This is the first diagnostic step because it confirms whether the bottleneck is on the delivery side or the processing side without instrumenting any application code.

Why this answer

Cloud Monitoring (Metrics Explorer) can show subscription backlog, Cloud Logging can show subscriber errors, and checking the subscription's backlog via gcloud or console helps assess the issue. Cloud Trace is for HTTP-based services, not Pub/Sub directly. Cloud Debugger is for code debugging, not Pub/Sub monitoring.

31
MCQhard

You need to drain a GKE node for maintenance. The node runs DaemonSet pods and pods using emptyDir volumes. Which kubectl drain command correctly handles these pods without causing disruption to critical system components?

A.kubectl drain NODE --force
B.kubectl drain NODE --ignore-daemonsets
C.kubectl drain NODE --delete-emptydir-data
D.kubectl drain NODE --ignore-daemonsets --delete-emptydir-data
AnswerD

Draining with `--ignore-daemonsets` leaves DaemonSet pods running, preserving node-level system components, while `--delete-emptydir-data` permits eviction of pods holding emptyDir volumes that would otherwise block the drain. Both flags together satisfy the stem's requirement to handle these pods without disrupting critical components.

Why this answer

The kubectl drain command evicts pods from a node. To handle DaemonSet pods (which should be ignored because they are managed by the DaemonSet controller) and pods with emptyDir volumes (which may need --delete-emptydir-data to force eviction), you use --ignore-daemonsets and --delete-emptydir-data.

32
Multi-Selectmedium

You need to export all logs from Cloud Logging to a BigQuery dataset for long-term analysis. The export should include logs from all projects in the organization. Which TWO actions should you take? (Choose two.)

Select 2 answers
A.Create a log sink with destination type Cloud Storage bucket
B.Create a log sink with destination type BigQuery dataset
C.Create a log-based metric to filter the logs
D.Create the sink at the organization level
E.Create the sink at the project level for each project individually
AnswersB, D

Creating a log sink with a BigQuery dataset destination routes matching log entries into BigQuery tables, satisfying the organisation-wide export requirement. Sink inclusion filters can scope the sink across all projects when created at the organisation level, enabling centralised long-term analysis without per-project configuration.

Why this answer

Option B is correct because a log sink's destination must be set to a BigQuery dataset to route log entries into BigQuery for long-term analysis; the sink uses the BigQuery destination and requires a dataset to exist. Option D is correct because creating the sink at the organization level allows it to aggregate and export logs from all projects in the organization in a single sink, which matches the requirement to include logs from all projects. Option A is incorrect because a Cloud Storage bucket destination exports logs to Cloud Storage, not BigQuery.

Option C is incorrect because log-based metrics create metric data for monitoring/alerting, not raw log export to BigQuery. Option E is incorrect because project-level sinks would need to be created per project and would not provide a single organization-wide export as required.

Exam trap

ACE often tests the difference between organization-level and project-level sinks; candidates may incorrectly choose project-level sinks for organization-wide log export, not realizing that organization-level sinks automatically cover all current and future projects.

33
MCQmedium

An engineer needs to attach an existing persistent disk to a Compute Engine instance. They have created the disk using 'gcloud compute disks create'. Which command should they use to attach it?

A.gcloud compute disks resize
B.gcloud compute instances attach-disk
C.gcloud compute instances add-disk
D.gcloud compute disks attach
AnswerB

gcloud compute instances attach-disk is the correct command: it attaches an existing zonal or regional persistent disk to a specified Compute Engine instance, using --disk and optionally --device-name. It works on both running and stopped instances, and it ensures the disk becomes visible as a block device in the instance's guest OS.

Why this answer

'gcloud compute instances attach-disk' attaches a disk to an instance. 'gcloud compute disks attach' does not exist. 'gcloud compute instances add-disk' is not a valid command. 'gcloud compute disks resize' resizes the disk.

34
Multi-Selecthard

You are responsible for monitoring a set of Compute Engine instances that run a critical web application. You want to be alerted when the average CPU utilization across all instances exceeds 80% for more than 5 minutes. You also want to receive a notification via email and SMS. Which TWO actions should you take? (Choose two.)

Select 2 answers
A.Configure a notification channel for email and SMS in Cloud Monitoring and attach it to the alerting policy.
B.Create a log-based alert in Cloud Logging that triggers when CPU utilization exceeds 80%.
C.Create a Cloud Monitoring alerting policy with a condition on the CPU utilization metric, setting the threshold to 80% and the duration to 5 minutes.
D.Install the Cloud Monitoring agent on each instance to collect CPU utilization metrics.
E.Set up an uptime check to monitor the CPU utilization of the instances.
AnswersA, C

This is correct because Cloud Monitoring supports notification channels for email, SMS, and other services. To receive alerts via email and SMS, you must create those channels and associate them with the alerting policy. Without notification channels, the alert would trigger but no notifications would be sent.

Why this answer

To alert on CPU utilization, you need a Cloud Monitoring alerting policy with a condition on the CPU metric, and you must attach notification channels for email and SMS. The Monitoring agent is not required for CPU metrics, log-based alerts are for logs, and uptime checks are for availability, not resource metrics.

Exam trap

The trap here is thinking that the Cloud Monitoring agent is needed for CPU metrics, when CPU is a built-in metric; also confusing log-based alerts with metric alerts.

35
MCQeasy

You want to export a subset of Cloud Logging logs to BigQuery for long-term analysis. Which method should you use?

A.Create a log-based metric and export the metric to BigQuery
B.Create a log sink with a filter and destination BigQuery
C.Set up a Cloud Function that triggers on logs and inserts into BigQuery
D.Use gcloud logging read and pipe to bq load
AnswerB

A log sink with a filter and a BigQuery destination is the fully managed, native way to export logs: Cloud Logging continuously routes any newly ingested log entries that match the filter into a specified BigQuery dataset. The sink automatically creates a table with the log schema, and you can use the _PARTITIONTIME pseudo-column for time-based partitioning. This gives reliable, near-real-time export without custom code or manual intervention.

Why this answer

A log sink is the native Cloud Logging mechanism for routing log entries to supported destinations, and BigQuery is a first-class sink destination. By attaching an inclusion filter to the sink, you can export only the subset of logs you care about, and Cloud Logging handles the delivery and schema management automatically. This is the designed, serverless, and most operationally sound approach for long-term log analysis in BigQuery.

Exam trap

The trap here is confusing log-based metrics (numeric Monitoring time series) with log sinks (raw log routing), causing candidates to pick the metric option when the question asks for exporting actual log data.

How to eliminate wrong answers

Option A is wrong because log-based metrics only produce numeric time-series counters/distributions in Cloud Monitoring; they do not carry the original log payload and cannot be exported as raw log rows to BigQuery. Option C is wrong because a Cloud Function triggered on logs is a custom, brittle workaround that duplicates what a native sink does, adds latency and cost, and is not the supported export path. Option D is wrong because 'gcloud logging read' piped to 'bq load' is a manual, batch, one-off operation that does not provide continuous, filtered, near-real-time export and requires you to manage schema and scheduling yourself.

36
MCQmedium

You want to monitor the uptime of an external HTTP endpoint every minute and receive an email notification if the endpoint is unavailable for more than two consecutive checks. What should you do?

A.Create a log-based alert in Cloud Logging that triggers on network errors
B.Create an uptime check in Cloud Monitoring, then create an alerting policy with condition 'metric threshold' for 'check_failed' and set notification channel to email
C.Use Cloud Functions to periodically call the endpoint and send an email on failure
D.Configure a TCP health check on the load balancer
AnswerB

Uptime checks in Cloud Monitoring are the managed, intended way to verify that an external HTTP endpoint is reachable and returning expected responses from multiple locations across the globe. The check_failed metric increments each time a probe fails, and a metric-threshold alerting policy lets you define a condition—for instance, when the number of failed checks is consistently above zero over a specified period—and route it to an email notification channel. This directly implements the requirement without custom code.

Why this answer

Cloud Monitoring uptime checks are purpose-built to probe external HTTP/HTTPS/TCP endpoints on a schedule (as frequent as once per minute), and alerting policies can trigger on the 'check_failed' metric with a threshold and duration that maps to 'two consecutive failures.' Email notification channels are natively supported, making this the correct, managed solution without custom code.

Exam trap

ACE often tests the difference between reactive log-based alerting and proactive uptime monitoring — candidates choose Cloud Logging alerts when the requirement is active endpoint probing.

How to eliminate wrong answers

Option A is wrong because log-based alerts in Cloud Logging react to log entries, not active probing of an external endpoint — there is no log generated if the endpoint simply goes down. Option C is wrong because Cloud Functions would require custom code, scheduling, state tracking for consecutive failures, and email integration — reinventing what Cloud Monitoring provides natively and adding operational overhead. Option D is wrong because a load balancer TCP health check only monitors backends behind that load balancer and does not send email alerts on its own; it also doesn't probe arbitrary external endpoints.

37
MCQmedium

You need to export logs from Cloud Logging to a BigQuery dataset for long-term analysis. What should you create?

A.An alerting policy with a log-based trigger
B.A log-based metric
C.An export job in BigQuery
D.A log sink with BigQuery as the destination
AnswerD

A log sink with BigQuery as the destination is the correct method: Cloud Logging's log router matches your chosen log entries and delivers them to a BigQuery dataset, where each daily collection becomes a table. You configure the destination by providing a dataset name, and the sink automatically handles batching and streaming writes. This is the officially supported, commonly used way to export logs to BigQuery for analytics.

Why this answer

To export logs from Cloud Logging to BigQuery, you create a log sink with BigQuery as the destination. Log sinks are the mechanism in Google Cloud for routing log entries to supported destinations, including BigQuery, Cloud Storage, and Pub/Sub. This allows for long-term storage and analysis.

Exam trap

ACE often tests the confusion between log-based metrics and log sinks; candidates may choose log-based metrics thinking they export logs, but metrics only aggregate data for monitoring, not export raw logs.

How to eliminate wrong answers

Option A is wrong because an alerting policy with a log-based trigger is used to notify when specific log events occur, not to export logs. Option B is wrong because a log-based metric counts or extracts values from logs for monitoring, but does not export the raw logs. Option C is wrong because an export job in BigQuery is not a native Cloud Logging feature; you cannot directly create an export job from Cloud Logging to BigQuery without a sink.

38
MCQhard

You need to drain a GKE node for maintenance, ensuring that daemonsets and pods using emptyDir volumes are handled properly. Which command should you use?

A.kubectl taint nodes NODE key=value:NoSchedule
B.kubectl drain NODE --ignore-daemonsets --delete-emptydir-data
C.kubectl delete node NODE
D.kubectl cordon NODE && kubectl delete pods --all
AnswerB

`kubectl drain` gracefully evicts all pods from the node while respecting PodDisruptionBudgets, making the node unschedulable and empty for maintenance. The `--ignore-daemonsets` flag skips DaemonSet-managed pods, which are intended to run on every node and would otherwise block eviction, while `--delete-emptydir-data` allows deletion of pods using emptyDir volumes, which would otherwise prevent the drain from finishing. These flags together ensure the command completes cleanly on nodes with these pod types.

Why this answer

kubectl drain with flags ignores daemonsets and deletes emptyDir pods.

39
MCQmedium

You need to resize a Compute Engine instance from n1-standard-4 to n1-highmem-8. The instance has a local SSD attached. What must you do before changing the machine type?

A.Stop the instance, change the machine type, then start the instance
B.Take a snapshot of the local SSD
C.Change the machine type without stopping
D.Detach the local SSD
AnswerA

To change the machine type of a Compute Engine instance, you must first stop it, which brings it to the TERMINATED state. While stopped, the persistent disks and instance settings remain intact, but any data on local SSDs is permanently lost because local SSDs are ephemeral storage tied to the host server. After updating the machine type, you start the instance; this process is the only supported way to resize an instance's vCPU and memory.

Why this answer

To change the machine type, the instance must be stopped. Local SSDs preserve data only if the instance is not stopped or terminated; however, when you stop the instance, local SSD data is lost. The correct procedure is to stop the instance, change the machine type, and then start it.

Data on local SSDs will be lost.

40
MCQhard

An application is experiencing intermittent high latency. Using Cloud Trace, an engineer identifies that the bottleneck is a Pub/Sub subscription with a large backlog. Which action would MOST directly help reduce the backlog?

A.Increase the ack deadline
B.Increase the maximum message size
C.Increase the message retention duration
D.Increase the number of subscribers
AnswerD

Increasing the number of subscribers (i.e., scaling out the subscriber fleet) directly raises the aggregate processing throughput of the subscription. Because the intermittent high latency is likely due to a backlog of messages accumulating faster than the current subscribers can drain, adding more subscribers allows messages to be pulled and processed in parallel, reducing the queue depth and lowering end-to-end latency. This is the correct scaling action for a latency problem caused by insufficient compute, assuming the subscribers are stateless and can process messages independently.

Why this answer

Increasing the number of subscribers (e.g., scaling out the subscriber application) will increase the processing rate and reduce backlog. Increasing the retention duration keeps messages longer, not reducing backlog. The ack deadline and message size are not the primary causes of backlog.

41
MCQmedium

You need to export all Cloud Logging logs from your project to BigQuery for long-term analysis. What should you create?

A.A Cloud Monitoring dashboard
B.A VPC flow log
C.A log-based alert
D.A log sink with destination BigQuery
AnswerD

A log sink in Cloud Logging's Router exports matching log entries to a destination such as BigQuery, Cloud Storage, or Pub/Sub. By configuring a sink with BigQuery as the destination and a filter that matches all logs (or empty filter), you continuously export your project's logs to a BigQuery dataset for analysis and long-term retention. This is the standard and only fully supported mechanism for exporting Cloud Logging logs to external services.

Why this answer

Log sinks route logs to supported destinations including BigQuery.

42
MCQmedium

You are investigating high latency in your application deployed on Compute Engine. You suspect a specific API call is taking longer than expected. Which Google Cloud tool should you use to analyze the latency of individual requests?

A.Cloud Debugger
B.Cloud Trace
C.Cloud Monitoring dashboards
D.Cloud Logging log explorer
AnswerB

Cloud Trace is a distributed tracing service designed to collect latency data from Google Cloud and measure time spent in each service and API call during a request. It provides detailed per-request traces with spans that show the timing of each operation, making it the correct tool to investigate high application latency. By analyzing the waterfall view of spans, you can identify the exact component responsible for the delay across distributed services.

Why this answer

Cloud Trace provides distributed tracing, allowing you to see the latency of individual requests and identify bottlenecks. It captures trace spans from supported frameworks and services.

43
MCQmedium

You manage a Google Kubernetes Engine (GKE) cluster and need to update the deployment 'web-app' to use a new container image tag 'v2'. You also want to ensure the update proceeds and, if it fails, roll back to the previous revision. Which set of commands should you use?

A.gcloud container clusters upgrade; kubectl rollout status; kubectl rollout undo
B.kubectl set image deployment/web-app web-app=gcr.io/myproject/web-app:v2; kubectl rollout status; kubectl rollout undo
C.kubectl edit deployment web-app; kubectl rollout status; kubectl delete deployment web-app
D.kubectl apply -f web-app.yaml; kubectl rollout status; kubectl rollout undo
AnswerB

kubectl set image directly updates the Deployment's pod template to reference the v2 image, which triggers a rolling update orchestrated by the Deployment controller. kubectl rollout status then watches that update to completion, returning a non-zero exit code if the rollout fails (e.g., due to crash-loops or insufficient readiness), which is the correct signal for conditional rollback. kubectl rollout undo reverts to the previous revision, but note it should be gated on that failure in practice; even so, this is the only option that uses the proper Kubernetes-native commands for image update, rollout monitoring, and rollback.

Why this answer

The correct command to update a deployment's container image is 'kubectl set image deployment/web-app web-app=gcr.io/myproject/web-app:v2', which directly updates the image tag for the specified container in the deployment. Following that, 'kubectl rollout status' monitors the update progress, and 'kubectl rollout undo' reverts to the previous revision if the update fails. This sequence ensures a controlled update with rollback capability.

Exam trap

The trap here is confusing cluster-level upgrades (gcloud container clusters upgrade) with application-level updates, or thinking that editing the deployment manually is equivalent to a controlled image update with rollback.

How to eliminate wrong answers

Option A is wrong because 'gcloud container clusters upgrade' upgrades the GKE cluster's Kubernetes version, not the application deployment image. Option C is wrong because 'kubectl edit deployment web-app' opens an editor for manual changes, which is error-prone and not scriptable, and 'kubectl delete deployment web-app' deletes the deployment entirely rather than rolling back. Option D is wrong because 'kubectl apply -f web-app.yaml' requires an updated manifest file and does not specifically target an image tag change; it could apply other unintended changes.

44
MCQmedium

You have a GKE cluster with a node pool that needs to scale automatically based on load. The cluster was created with autoscaling disabled. Which command enables autoscaling on an existing node pool?

A.gcloud container node-pools create my-pool --enable-autoscaling
B.kubectl autoscale node-pool my-pool --min=1 --max=10
C.gcloud container node-pools update my-pool --cluster=my-cluster --enable-autoscaling --min-nodes=1 --max-nodes=10
D.gcloud container clusters update my-cluster --enable-autoscaling
AnswerC

This is the correct command because it targets an existing node pool (`my-pool`) within the specified cluster and toggles the GKE cluster autoscaler on for that pool. The `--min-nodes=1` and `--max-nodes=10` flags define the scaling boundaries, allowing the pool to resize within those limits based on resource demand. The `--cluster` flag scopes the operation to the right cluster, and the update command modifies the live pool without recreating it.

Why this answer

The correct command is `gcloud container node-pools update` with `--enable-autoscaling` and min/max node flags, because autoscaling is a property of the node pool, not the cluster, and it must be enabled on an existing pool via the update verb. The `--cluster` flag identifies the parent cluster, and `--min-nodes`/`--max-nodes` define the scaling bounds. This is the only option that both targets an existing node pool and uses the correct gcloud subcommand.

Exam trap

The trap here is confusing cluster-level and node-pool-level operations — candidates often pick the cluster update command or the create command, forgetting that autoscaling is configured per node pool and requires the `node-pools update` verb.

How to eliminate wrong answers

Option A is wrong because `gcloud container node-pools create` creates a brand-new node pool rather than enabling autoscaling on the existing one, which would leave the original pool unchanged and add unnecessary resources. Option B is wrong because `kubectl autoscale` operates on Kubernetes workload resources (Deployments, ReplicaSets, StatefulSets) via HorizontalPodAutoscaler, not on GKE node pools; there is no `node-pool` resource type for kubectl autoscale. Option D is wrong because `gcloud container clusters update --enable-autoscaling` is not a valid way to enable node pool autoscaling — cluster-level update commands do not carry the node pool's min/max node configuration, and the flag belongs to node-pool operations.

45
Multi-Selectmedium

You need to set up log-based alerting in Cloud Logging to send notifications when a specific error pattern appears in your application logs. Which TWO components are required to accomplish this?

Select 2 answers
A.An alerting policy
B.A log sink
C.A Cloud Pub/Sub topic
D.An uptime check
E.A log-based metric
AnswersA, E

The alerting policy is the actual alerting mechanism in Cloud Logging. It defines the conditions that trigger an incident, such as a threshold on a metric (e.g., a log-based metric exceeding a value) and specifies the notification channels (email, Slack, etc.) to receive alerts. Without an alerting policy, a log-based metric only counts or samples log entries; it does not perform any active monitoring or notify anyone.

Why this answer

Option A, an alerting policy, is required because it is the Cloud Monitoring resource that defines the condition to watch and the notification channels to fire when that condition is met, which is what actually sends the notifications. Option E, a log-based metric, is required because log-based alerting in Cloud Logging works by defining a metric that counts or extracts values from log entries matching the specified error pattern, and the alerting policy then evaluates that metric. Together, the log-based metric translates the log filter into a time series and the alerting policy triggers notifications when that series crosses the threshold.

Option B, a log sink, is not required because sinks route or export log entries to destinations such as Cloud Storage or BigQuery and are not part of the alerting evaluation path. Option C, a Cloud Pub/Sub topic, is not required because Pub/Sub is only needed for log routing/export or for notification delivery if explicitly chosen as a channel, not for creating the log-based alert itself. Option D, an uptime check, is not required because uptime checks probe endpoint availability over HTTP/TCP and are unrelated to matching error patterns in log entries.

Exam trap

ACE often tests the confusion between log routing (sinks) and log alerting — candidates who think a log sink or Pub/Sub topic is required for alerting miss that the essential pair is a log-based metric plus an alerting policy.

46
Multi-Selectmedium

You are troubleshooting a Pub/Sub subscription that is not delivering messages promptly. Which THREE factors should you investigate? (Choose THREE.)

Select 3 answers
A.The subscription's backlog size
B.The topic's retention duration
C.The subscriber's processing latency
D.The message ordering key
E.The acknowledgment deadline
AnswersA, C, E

The subscription's backlog size is the primary indicator of delivery problems: it counts messages that have been published but not yet acknowledged. When troubleshooting a subscription that is not delivering, an ever-growing backlog means messages are arriving faster than the subscriber can process them, or the subscriber has stopped pulling entirely. Large backlog also correlates with slow processing and can help you decide whether to scale out subscribers or inspect subscriber logs.

Why this answer

Common causes include backlog, subscriber latency, and ack deadlines.

47
MCQmedium

You have a Compute Engine VM instance that is currently running. You need to resize it to a different machine type. What must you do first?

A.Stop the instance, then use gcloud compute instances set-machine-type, then start the instance.
B.Use gcloud compute instances update --machine-type while the instance is running.
C.Detach all disks, change machine type, then reattach disks.
D.Create a snapshot of the disk and use it to create a new instance with the desired machine type.
AnswerA

Stopping the instance transitions it to the TERMINATED state, which releases the underlying host resources while preserving the boot disk, metadata, and attachment of persistent disks. The `gcloud compute instances set-machine-type` command can then change the vCPU and memory allocation, and after that you start the instance. This is the correct workflow because Compute Engine rejects machine type changes on running instances.

Why this answer

Changing the machine type requires the VM to be in a stopped state. You must stop the instance, change the machine type, then start it.

48
Multi-Selectmedium

A company wants to automate the response to specific log entries by triggering a Cloud Function. Which THREE components are required? (Choose 3)

Select 3 answers
A.Cloud Function (Pub/Sub trigger)
B.Cloud Logging log sink
C.Pub/Sub topic
D.BigQuery dataset
E.Cloud Monitoring notification channel
AnswersA, B, C

A Cloud Function with a Pub/Sub trigger is the compute piece that executes your custom response logic asynchronously. When a message lands on the subscribed topic, the function is invoked with the message payload, letting you parse the log data and call external APIs, send alerts, or modify resources. This event-driven model avoids maintaining a server and scales automatically with message volume. Without this function, the sink and topic would merely transport logs with no automated reaction.

Why this answer

Log entries must be routed to a Pub/Sub topic via a log sink. The Cloud Function subscribes to that topic (triggered by Pub/Sub). The log sink is the exporter, Pub/Sub is the intermediary, and Cloud Function is the action.

A notification channel is for alerts, not triggers. BigQuery is not needed.

49
MCQeasy

An engineer needs to monitor the external HTTP availability of a web application hosted on Compute Engine. Which Cloud Monitoring feature should they use?

A.Uptime check
B.Dashboard
C.Metric Explorer
D.Log-based alert
AnswerA

An uptime check is a Cloud Monitoring synthetic probe that periodically sends an HTTP(S) request to the specified URL from configurable global locations. It validates availability by checking for expected HTTP status codes, response time thresholds, and optional content matches, and it emits metrics such as uptime, latency, and check success. This is the correct choice because it actively measures external HTTP reachability from outside the network, which is exactly what is needed to monitor external availability.

Why this answer

Uptime checks are designed to verify that a resource is accessible and measure response latency from various locations. They can check HTTP/HTTPS/TCP endpoints.

50
MCQmedium

A Cloud Run service is experiencing high latency. You suspect one revision is causing the issue. The service is configured to split traffic 90% to revision A and 10% to revision B. You want to gradually shift traffic back to revision A only. Which command should you use?

A.kubectl set traffic my-service --revision=my-service-00001=100
B.gcloud run services update-traffic my-service --to-revisions=my-service-00001=100
C.gcloud run revisions delete my-service-00002
D.gcloud run services update my-service --set-revision my-service-00001
AnswerB

This is the correct, supported command for adjusting traffic on a Cloud Run service. The 'update-traffic' subcommand directly modifies the revision routing percents, and '--to-revisions' allows explicit targeting of a specific revision; here, setting 'my-service-00001=100' routes all live traffic to the known-good revision A. This immediately reduces load on the suspect revision B and is exactly how you roll back a bad deployment on Cloud Run.

Why this answer

The `gcloud run services update-traffic` command is the correct tool for adjusting traffic distribution across Cloud Run revisions. Using `--to-revisions=my-service-00001=100` assigns 100% of traffic to revision A (my-service-00001), effectively removing revision B from the serving path. This is the supported, declarative way to shift traffic on a Cloud Run service without redeploying.

Exam trap

The trap here is confusing kubectl syntax with gcloud syntax — candidates who work with Kubernetes may instinctively pick the kubectl option, but Cloud Run traffic management is exclusively a gcloud/API operation.

How to eliminate wrong answers

Option A is wrong because `kubectl set traffic` is not a valid kubectl subcommand — Cloud Run traffic management is done through gcloud or the Cloud Run Admin API, not kubectl. Option C is wrong because deleting revision B does not shift traffic to revision A; it only removes the revision, and traffic would still need to be reassigned explicitly. Option D is wrong because `gcloud run services update --set-revision` is not a valid flag; the correct flag for traffic assignment is `--to-revisions` under `update-traffic`.

51
MCQeasy

You need to store application logs from a Compute Engine instance in a way that allows you to search and analyze them later. The logs should be retained for 30 days. Which Google Cloud service should you use?

A.BigQuery
B.Cloud Logging
C.Cloud Storage
D.Cloud Monitoring
AnswerB

Cloud Logging is a fully managed service that ingests, stores, and allows you to search and analyze logs from Google Cloud resources, including Compute Engine instances. By default, logs are retained for 30 days, which matches the requirement. You can also create log sinks to export logs for longer retention if needed, but for the stated 30-day period, Cloud Logging alone is sufficient.

Why this answer

Cloud Logging is the native Google Cloud service for log ingestion, storage, and analysis. It automatically collects logs from Compute Engine instances and other resources, retains them for 30 days by default, and provides a powerful query language for searching. Other services like Cloud Storage or BigQuery can be used for export, but they are not designed for direct log management.

Exam trap

The trap here is confusing Cloud Monitoring with Cloud Logging; Monitoring is for metrics and alerting, while Logging is specifically for logs.

52
Multi-Selecthard

You are deploying a new version of a microservice to Google Kubernetes Engine (GKE). You want to minimize downtime and ensure that traffic is only routed to pods that are ready to serve requests. Which TWO actions should you take? (Choose two.)

Select 2 answers
A.Configure a readiness probe for the pods.
B.Configure a liveness probe for the pods.
C.Expose the deployment using a LoadBalancer service.
D.Set the pod's restartPolicy to Always.
E.Use a rolling update strategy with maxSurge and maxUnavailable.
AnswersA, E

A readiness probe determines whether a pod is ready to accept traffic. When a pod fails the readiness probe, it is removed from the Service's endpoints, preventing traffic from being routed to it. This ensures that only healthy pods receive requests, which is essential for zero-downtime deployments and maintaining service availability during updates.

Why this answer

To minimize downtime during a deployment, you need both a readiness probe and a rolling update strategy. Readiness probes ensure that only pods that are ready to serve traffic are added to the Service's endpoints, while rolling updates gradually replace old pods with new ones, maintaining availability. Together, they prevent traffic from reaching unready pods and ensure a smooth transition.

Exam trap

The trap here is focusing on liveness probes or restart policies, which handle container restarts but do not control traffic routing during deployments.

53
Multi-Selecthard

An engineer is troubleshooting a Compute Engine instance that is unreachable via SSH. They suspect a firewall rule is blocking traffic. Which TWO actions should they take to diagnose the issue? (Choose 2)

Select 2 answers
A.Create a Cloud Monitoring alert for packet loss
B.View Cloud Logging for firewall rule logs
C.Run gcloud compute ssh --dry-run
D.Use Cloud Trace to analyze network latency
E.Check VPC firewall rules in Cloud Console
AnswersB, E

Viewing Cloud Logging for firewall rule logs is the direct way to see whether VPC firewall rules are dropping or allowing traffic. Firewall rule logging records each connection attempt with details like source IP, destination IP, port, protocol, and the action (allow or deny). If the Compute Engine instance is unreachable due to a firewall rule, these logs will show the denied packets, making this a reliable troubleshooting step.

Why this answer

In Cloud Logging, you can view firewall logs (if VPC flow logs are enabled, but firewall rules logging can be enabled per rule). Checking VPC firewall rules in the Cloud Console allows you to verify the rules. Cloud Trace is for latency, Cloud Monitoring for metrics, and gcloud compute ssh is for connecting, not diagnosing firewall rules.

54
MCQhard

You need to perform a rolling update of a GKE deployment and ensure that during the update, the new pods are ready before terminating the old ones. You have already set the update strategy to RollingUpdate. Which kubectl command sequence should you use to update the image and monitor the rollout?

A.gcloud container clusters upgrade my-cluster; kubectl get deployments
B.kubectl set image deployment/myapp myapp=gcr.io/myproject/myapp:v2; kubectl rollout status deployment/myapp
C.kubectl edit deployment myapp; kubectl get pods; kubectl delete pod old-pod
D.kubectl apply -f deployment.yaml; kubectl rollout undo deployment/myapp
AnswerB

kubectl set image updates the Deployment's pod template to gcr.io/myproject/myapp:v2, which triggers the Deployment controller to create a new ReplicaSet and incrementally replace old pods while respecting maxSurge/maxUnavailable. kubectl rollout status then blocks until the new ReplicaSet becomes ready and the old ReplicaSet is scaled down, confirming the rolling update completed successfully. This is the standard declarative workflow for updating an app version.

Why this answer

`kubectl set image` updates the container image on the deployment, which triggers the RollingUpdate strategy to create new pods and only terminate old ones once the new pods pass readiness probes. `kubectl rollout status` then blocks and reports the progress of that rollout until it completes or fails. Together they satisfy both the update and monitoring requirements.

Exam trap

ACE often tests the difference between updating a workload (`set image`) and rolling back (`rollout undo`), so candidates who see 'rollout' and grab the undo command pick the wrong answer.

How to eliminate wrong answers

Option A is wrong because `gcloud container clusters upgrade` upgrades the GKE control plane or node version, not a workload's container image, and `kubectl get deployments` gives only a snapshot, not rollout progress. Option C is wrong because manually editing and deleting pods bypasses the controlled RollingUpdate and risks downtime; it also does not monitor rollout status. Option D is wrong because `kubectl rollout undo` rolls back to a previous revision, which is the opposite of performing the update.

55
Multi-Selecteasy

You need to set up an alerting policy to notify your team via email and Slack when a Compute Engine instance's CPU utilization exceeds 80% for 5 minutes. Which two resources must you configure? (Choose two.)

Select 2 answers
A.A Cloud Function to check CPU and send Slack message
B.A metric threshold condition on the 'compute.googleapis.com/instance/cpu/utilization' metric
C.An uptime check for the external IP of the instance
D.A notification channel of type 'email'
E.A log-based alert for the 'compute.googleapis.com/instance' log
AnswersB, D

The correct condition uses a metric threshold on the time series compute.googleapis.com/instance/cpu/utilization. This metric is emitted automatically from GCE instances and can be queried with a threshold (e.g., > 80%) aligned over a defined period such as 5 minutes. When the condition's duration (e.g., 'for 5 minutes') is met, the alerting policy enters the firing state and notifies any attached channels. This is the native, fully integrated way to alert on CPU load.

Why this answer

Option B is correct because a Cloud Monitoring alerting policy requires an alerting condition, and a metric threshold condition on 'compute.googleapis.com/instance/cpu/utilization' with a threshold of 80% and a duration of 5 minutes directly expresses the required trigger. Option D is correct because notification channels define where alerts are delivered; an 'email' notification channel is needed to notify the team via email (a Slack channel would be configured as an additional notification channel of type Slack). Option A is not required because Cloud Monitoring natively evaluates metrics and sends notifications without a Cloud Function.

Option C is not relevant because uptime checks test endpoint availability, not CPU utilization. Option E is not relevant because log-based alerts trigger on log entries, not on a CPU metric threshold.

Exam trap

ACE often tests the misconception that additional compute resources (like Cloud Functions) are needed for alerting, when Cloud Monitoring natively handles metric-based alerts.

56
MCQhard

Your GKE cluster is running a deployment with a container image my-app:v1. You need to update it to my-app:v2 and monitor the rollout progress. Which commands should you use?

A.gcloud compute instances update-container and kubectl get events
B.kubectl edit deployment/my-app and change the image, then kubectl rollout undo if needed
C.kubectl set image deployment/my-app my-app=my-app:v2 followed by kubectl rollout status deployment/my-app
D.gcloud container clusters upgrade and kubectl get pods
AnswerC

kubectl set image deployment/my-app my-app=my-app:v2 imperatively updates the container image of the specified container in the Deployment, which immediately triggers a new ReplicaSet and rolling update. kubectl rollout status deployment/my-app then blocks and reports the status of that rollout until it completes, satisfying the requirement to update and monitor progress in one straightforward command sequence.

Why this answer

The correct approach uses `kubectl set image` to declaratively update the container image in the deployment, which triggers a rolling update, followed by `kubectl rollout status` to watch the rollout progress until completion. This is the standard Kubernetes-native workflow for image updates and monitoring.

Exam trap

The trap here is confusing GKE workload management with Compute Engine container management, or assuming `kubectl edit` is the preferred method for image updates when the exam expects the imperative `kubectl set image` + `kubectl rollout status` pattern.

How to eliminate wrong answers

Option A is wrong because `gcloud compute instances update-container` targets Compute Engine VMs running containers (via the container-vm or COS), not GKE deployments, and `kubectl get events` only shows cluster events, not rollout progress. Option B is wrong because while `kubectl edit` can change the image, it is an interactive manual edit rather than the recommended imperative command, and `kubectl rollout undo` is only used to roll back, not to monitor progress. Option D is wrong because `gcloud container clusters upgrade` upgrades the cluster's Kubernetes version, not a workload's image, and `kubectl get pods` shows pod status but not a structured rollout status.

57
MCQeasy

You need to alert when the CPU utilization of your Compute Engine instance exceeds 80% for 5 minutes. What should you create in Cloud Monitoring?

A.An uptime check
B.A metric threshold alerting policy
C.A log-based alert
D.A dashboard chart
AnswerB

In Cloud Monitoring, you create an alerting policy with a condition that uses a threshold for a metric such as 'compute.googleapis.com/instance/cpu/utilization'. The policy samples the metric stream over an alignment period and triggers when the value (e.g., average CPU utilization) crosses the threshold for a specified duration. This is exactly the native mechanism for CPU utilization alerts.

Why this answer

Alerting on a sustained CPU utilization threshold (over 80% for 5 minutes) requires a metric threshold alerting policy in Cloud Monitoring. This policy evaluates the CPU utilization metric against a threshold over a defined duration and triggers notifications when the condition is met. It is the standard mechanism for resource-based alerting.

Exam trap

The trap is confusing alerting mechanisms — candidates may pick uptime checks or log-based alerts because they sound like monitoring, but only a metric threshold policy evaluates numeric resource metrics over time.

How to eliminate wrong answers

Option A is wrong because an uptime check monitors endpoint availability (HTTP/TCP responses), not CPU utilization. Option C is wrong because a log-based alert triggers on log entries matching a filter, not on metric values like CPU percentage. Option D is wrong because a dashboard chart is a visualization tool — it displays metrics but does not evaluate conditions or send alerts.

58
MCQhard

You need to create a log-based metric that counts the number of 5xx errors from your application logs. The logs are in Cloud Logging and contain a field "httpRequest.status". Which filter should you use when creating the metric?

A.httpRequest.status:5*
B.severity=ERROR AND "5xx"
C.httpRequest.status = 500 OR httpRequest.status = 501 OR httpRequest.status = 502
D.httpRequest.status >= 500
AnswerD

This filter uses a comparison operator on the numeric field httpRequest.status. In Cloud Logging, filters support comparison operators like >= for numeric values, so this will match any log entry where the HTTP response status is 500 or higher, capturing all server error statuses (5xx). This is the recommended approach because it is concise and semantically correct.

Why this answer

Log-based metrics use Cloud Logging filter language to select log entries.

59
MCQmedium

You have a BigQuery table with billions of rows. You need to create a new table with the same schema and copy all data from the original table. Which approach is most efficient?

A.Use bq load with an empty file to create the table, then insert data row by row.
B.Export the original table to Cloud Storage as Avro, then load into the new table.
C.Use bq query --destination_table mydataset.newtable 'SELECT * FROM mydataset.original'
D.Use bq cp (copy) command.
AnswerD

bq cp performs a server-side metadata copy, duplicating the schema and all rows without streaming data through the client, so billions of rows transfer in seconds and avoid the cost and time of a query-based rewrite.

Why this answer

The bq cp command performs a server-side copy of a table within BigQuery, duplicating both schema and data without exporting or re-importing. It is the fastest and most cost-effective method because no data leaves BigQuery's storage layer. This is the canonical approach for cloning large tables.

Exam trap

ACE often tests the misconception that SELECT * with a destination table is equivalent to a copy, ignoring that it triggers a full scan and query charges.

How to eliminate wrong answers

Option A is wrong because row-by-row inserts are extremely slow and expensive at billions-of-rows scale, and bq load with an empty file does not create a populated table. Option B is wrong because exporting to Cloud Storage and reloading incurs egress/storage costs and is far slower than a native copy. Option C is wrong because a SELECT * query with a destination table scans all data, consuming query slots and bytes billed, whereas bq cp is metadata-level and free.

60
MCQeasy

A site reliability engineer needs to be notified immediately when the error rate of a production microservice exceeds 5% over a 5-minute window. Which type of alerting policy should be used?

A.Uptime check alert
B.Pub/Sub notification hook
C.Metric threshold alert
D.Log-based alert (log metric trigger)
AnswerC

A metric threshold alert continuously evaluates a time-series metric (e.g., error rate, request latency, or CPU utilization) against a user-defined threshold, such as 'error rate > 5% for 5 minutes', and immediately triggers a notification when the condition is met. This is precisely the right tool for an SRE who needs to be notified when application error rates exceed an acceptable level, because it supports real-time aggregation, sliding windows, and alerting policies with multiple notification channels. It is the correct answer.

Why this answer

A metric threshold alert is the correct choice because it triggers when a specific metric (error rate) crosses a defined threshold (5%) over a specified time window (5 minutes). This type of alerting policy is designed for monitoring numeric metrics and firing alerts based on conditions, making it ideal for this requirement.

Exam trap

ACE often tests the difference between alerting policy types, and candidates may confuse log-based alerts with metric threshold alerts, especially when the metric is derived from logs.

How to eliminate wrong answers

Option A is wrong because an uptime check alert monitors the availability of a service (e.g., HTTP 200 response) and does not evaluate error rates. Option B is wrong because a Pub/Sub notification hook is a mechanism to send alerts (e.g., to Pub/Sub), not a policy type that defines the condition. Option D is wrong because a log-based alert triggers on the occurrence of specific log entries, not on aggregated metrics like error rate; while you could create a log-based metric, the alerting policy itself would still be a metric threshold alert.

61
MCQmedium

You are configuring an uptime check for an HTTPS endpoint that returns a JSON response. The check should validate that the response contains a specific field "status":"ok". Which uptime check option should you use?

A.Enable SSL hostname verification
B.Configure a notification channel
C.Add a content match with a regular expression
D.Create a log-based alert for the endpoint
AnswerC

Adding a content match with a regular expression is the direct way to verify a specific string pattern in the HTTPS response body. Cloud Monitoring uptime checks accept both substring and regex content matches, allowing you to assert that the page contains a particular marker or dynamic token. This confirms the endpoint is serving expected application content, not just a reachable server.

Why this answer

Uptime checks can validate response content using content matching.

62
Multi-Selecthard

Your GKE cluster is running an older version of Kubernetes. You need to upgrade the cluster's control plane and node pools. Which two steps should you perform? (Choose two.)

Select 2 answers
A.Create a new cluster with the desired version and migrate workloads
B.Drain all nodes using kubectl drain before upgrading
C.Manually update the kubelet version on each node
D.Upgrade the cluster's control plane using gcloud container clusters upgrade
E.Upgrade node pools using gcloud container node-pools upgrade
AnswersD, E

Upgrading the cluster's control plane with `gcloud container clusters upgrade` is the correct first step because GKE enforces a maximum version skew between the control plane and node pools—typically one minor version. The control plane must be on the target version before node pools can be upgraded, and this command without a `--node-pool` flag updates only the control plane. This ensures the Kubernetes API server and scheduler are consistent with the target version, reducing the risk of API deprecations or incompatibility. It is the only supported way to perform an in-place control plane upgrade while preserving cluster identity and state.

Why this answer

Option D is correct because the control plane of a GKE cluster is upgraded with the gcloud container clusters upgrade command, which targets the cluster's master components to the desired Kubernetes version. Option E is correct because node pools are upgraded separately using gcloud container node-pools upgrade, which rolls out the new node version pool by pool while respecting surge and drain settings. These two steps match GKE's standard upgrade model: first the control plane, then the node pools.

Option A is not required because in-place upgrades are supported and recreating the cluster is unnecessary. Option B is not a required step because GKE handles node draining automatically during a node pool upgrade. Option C is incorrect because kubelet versions on GKE nodes are managed by the node pool upgrade process, not by manual per-node changes.

Exam trap

The trap is confusing workload operations (draining nodes, migrating clusters) with the actual GKE upgrade commands — candidates must know the two distinct gcloud commands for control plane and node pool upgrades.

63
MCQeasy

Your team uses Cloud Logging to store application logs. You want to create a metric that counts the number of ERROR log entries per service. Which type of log-based metric should you create?

A.Distribution metric
B.Boolean metric
C.Counter metric
D.Gauge metric
AnswerC

A counter metric is the correct log-based metric type for this use case, because it increments by one for every log entry that matches the specified filter, such as severity=ERROR. This gives the total number of error logs over the selected time window, which is exactly what the team wants to track. In Cloud Logging, you define a counter-based log metric with a filter and then use it in Monitoring charts or alerts.

Why this answer

A counter metric in Cloud Logging counts the number of log entries that match a filter. To count ERROR log entries per service, you create a log-based counter metric with a filter like severity=ERROR and a label extractor for the service name. Counter metrics are the correct type for tallying occurrences over time.

Exam trap

The trap is mixing up metric types — candidates often choose distribution or gauge because they sound more sophisticated, but only a counter metric is designed to count occurrences of matching log entries.

How to eliminate wrong answers

Option A is wrong because distribution metrics capture numeric values from log entries (e.g., latency) and produce histograms, not counts of matching entries. Option B is wrong because boolean metrics track whether a specific log entry occurred (true/false) and are used for presence/absence, not cumulative counts. Option D is wrong because gauge metrics record the latest numeric value from a log entry (e.g., current memory usage) and do not aggregate counts.

64
MCQhard

Your GKE cluster has a node pool that you want to enable autoscaling on. The initial node count is 3, and you want the cluster to scale between 1 and 10 nodes. Which command should you use?

A.gcloud container clusters update my-cluster --enable-autoscaling --min-nodes 1 --max-nodes 10 --region us-central1
B.gcloud container clusters update my-cluster --enable-autoscaling --min-size 1 --max-size 10 --region us-central1
C.gcloud container node-pools update my-pool --cluster=my-cluster --enable-autoscaling --min-nodes 1 --max-nodes 10 --region us-central1
D.gcloud container node-pools update my-pool --cluster=my-cluster --autoscaling --min 1 --max 10 --region us-central1
AnswerC

This is the correct command because autoscaling is a node-pool-level feature. It uses `gcloud container node-pools update`, points at the specific pool with `--cluster=my-cluster`, and enables the Cluster Autoscaler with the valid `--enable-autoscaling` flag. The `--min-nodes 1 --max-nodes 10` range constrains the pool size, and `--region us-central1` correctly specifies the regional control plane where the cluster lives.

Why this answer

The correct command to enable autoscaling on an existing node pool is 'gcloud container node-pools update my-pool --cluster=my-cluster --enable-autoscaling --min-nodes 1 --max-nodes 10 --region us-central1'. This command targets the node pool specifically and uses the correct flags for minimum and maximum node counts.

Exam trap

ACE often tests the exact gcloud command syntax for node pool autoscaling, and candidates might confuse cluster-level and node-pool-level commands or use incorrect flag names.

How to eliminate wrong answers

Option A is wrong because it attempts to enable autoscaling at the cluster level, which is not how node pool autoscaling is configured; the command 'gcloud container clusters update' with '--enable-autoscaling' is used for cluster autoscaler but requires node pool specification? Actually, for cluster autoscaler, you enable it per node pool. Option B is wrong because it uses '--min-size' and '--max-size' which are not valid flags for node pool autoscaling; the correct flags are '--min-nodes' and '--max-nodes'. Option D is wrong because it uses '--autoscaling' instead of '--enable-autoscaling' and '--min'/'--max' instead of '--min-nodes'/'--max-nodes'.

65
MCQeasy

You need to load a CSV file from Cloud Storage into an existing BigQuery table. Which bq command should you use?

A.bq query --source_format=CSV 'SELECT * FROM mydataset.mytable'
B.bq load --source_format=CSV mydataset.mytable gs://mybucket/myfile.csv
C.bq insert mydataset.mytable gs://mybucket/myfile.csv
D.bq import mydataset.mytable gs://mybucket/myfile.csv
AnswerB

bq load is the correct BigQuery CLI command to initiate a batch load job from Cloud Storage. It creates a load job that reads the CSV file at the given URI, parses it according to the specified --source_format, and writes rows into the target table (mydataset.mytable), which can be appended to or replace. This is the standard, idempotent way to bulk-load CSV data into BigQuery.

Why this answer

The bq load command loads data into a BigQuery table. You specify the source format (CSV) and the location of the file in Cloud Storage.

66
MCQmedium

You notice that a deployment in your GKE cluster is running an outdated image. You need to update the deployment to use the new image 'gcr.io/my-project/my-app:v2'. Which kubectl command should you use?

A.kubectl set image deployment/my-deployment my-app=gcr.io/my-project/my-app:v2
B.kubectl rollout restart deployment my-deployment --image gcr.io/my-project/my-app:v2
C.kubectl update deployment my-deployment --image gcr.io/my-project/my-app:v2
D.kubectl replace deployment my-deployment --image gcr.io/my-project/my-app:v2
AnswerA

kubectl set image deployment/my-deployment my-app=gcr.io/my-project/my-app:v2 is the correct imperative command to update a container image inside a Deployment. The container name (my-app) must exactly match the container name defined in the Deployment's pod spec, and the command updates the pod template so the Deployment controller creates a new ReplicaSet and performs a rolling update. This is the canonical kubectl syntax for changing an image without editing a manifest.

Why this answer

The correct command to update a deployment's container image is 'kubectl set image', which updates the image for a specific container within the deployment. The syntax 'deployment/my-deployment my-app=gcr.io/my-project/my-app:v2' specifies the deployment name and the container name with the new image. This triggers a rolling update.

Exam trap

The trap is confusing commands that sound similar: 'kubectl set image' is the correct one, but candidates might choose 'rollout restart' thinking it updates the image, or 'replace' thinking it's a general update command.

How to eliminate wrong answers

Option B is wrong because 'kubectl rollout restart' is used to restart a deployment (e.g., to pick up a ConfigMap change) but does not change the image; the '--image' flag is not valid for this command. Option C is wrong because 'kubectl update' is not a valid kubectl command; the correct verb is 'set image' or 'apply'. Option D is wrong because 'kubectl replace' requires a full manifest file and does not support an '--image' flag; it would replace the entire deployment configuration.

67
MCQhard

You need to drain a GKE node for maintenance. The node is running a DaemonSet and some pods with emptyDir volumes. Which kubectl command should you use to safely drain the node without causing errors?

A.kubectl drain node-name --force
B.kubectl drain node-name --ignore-daemonsets --delete-emptydir-data
C.kubectl drain node-name --ignore-daemonsets
D.kubectl cordon node-name && kubectl delete pods --all --grace-period=0
AnswerB

This is the correct drain command because --ignore-daemonsets tells kubectl to skip evicting Pods that are managed by DaemonSets (they would just be recreated on the same node), and --delete-emptydir-data lets the drain proceed even if Pods have emptyDir volumes that will be lost. The drain cordons the node, then gracefully evicts remaining workload Pods while honoring PodDisruptionBudgets, making it safe for planned maintenance. These flags are the standard pair used when a GKE node contains DaemonSets and emptyDir-backed pods.

Why this answer

The correct drain command must account for both DaemonSet-managed pods (which cannot be evicted and require --ignore-daemonsets) and pods using emptyDir volumes (which hold local data and require --delete-emptydir-data to allow eviction). Combining both flags lets kubectl drain the node without erroring on these two conditions.

Exam trap

ACE often tests whether candidates know that drain requires explicit flags for DaemonSets and emptyDir pods — candidates pick --ignore-daemonsets alone and are surprised when drain fails on emptyDir volumes.

How to eliminate wrong answers

Option A is wrong because --force alone bypasses unmanaged pods but does not handle DaemonSet pods or emptyDir data, so drain will still fail on those. Option C is wrong because --ignore-daemonsets handles DaemonSet pods but omits --delete-emptydir-data, causing drain to abort when it encounters pods with emptyDir volumes. Option D is wrong because cordon only marks the node unschedulable and does not evict pods; deleting pods with --grace-period=0 forcibly terminates them without respecting PodDisruptionBudgets or graceful shutdown, risking data loss and service disruption.

68
Multi-Selecteasy

You want to create a monitoring dashboard that shows a time-series chart of CPU utilization for a specific Compute Engine instance. Which THREE components do you need to configure? (Choose three.)

Select 3 answers
A.Choose a time aggregation function (e.g., mean, max)
B.Select the resource type: 'gce_instance' and filter by the instance ID
C.Create a log-based metric for CPU utilization
D.Select the metric: 'compute.googleapis.com/instance/cpu/utilization'
E.Set up a notification channel to send alerts
AnswersA, B, D

Choosing a time aggregation function is essential because raw metric samples arrive at irregular intervals and need to be aligned to a fixed time step for a coherent time series. The aggregation function (e.g., mean, max, sum) reduces multiple points within each alignment window into a single value, which determines the chart's shape and sensitivity to spikes. Without this step, the dashboard may render unusable, too-dense data or fail to produce a meaningful trend. For CPU utilization, 'mean' is typical for overall usage, while 'max' can highlight peak behavior.

Why this answer

In Cloud Monitoring, to create a chart you need to select a metric, a resource, and a time aggregation function.

69
Multi-Selecthard

Your Cloud Run service has a new revision that you want to gradually shift traffic to. You want to send 10% of traffic to the new revision and 90% to the current one. Which TWO steps are required? (Choose TWO.)

Select 2 answers
A.Set a new default URL for the new revision.
B.Delete the old revision.
C.Create the new revision by updating the service with a new image tag.
D.Enable VPC ingress for the new revision.
E.Use gcloud run services update-traffic to set traffic percentages.
AnswersC, E

Updating the service with a new image tag, for example via `gcloud run deploy`, is what creates a new revision. A revision cannot be manually created in isolation—it is always the result of deploying a new container image or configuration change. This step is a prerequisite because the later `update-traffic` command must reference the new revision's name to assign it a percentage of incoming requests.

Why this answer

You first create the new revision (by updating the service) and then modify traffic percentages.

70
MCQeasy

You have a Compute Engine VM that is running a critical application. You need to change its machine type from n1-standard-4 to n2-standard-8. What is the correct procedure?

A.Stop the instance, then use gcloud compute instances set-machine-type, then start the instance
B.Use gcloud compute instances update --machine-type n2-standard-8 while the instance is running
C.Delete the instance and create a new one with the desired machine type
D.Use gcloud compute instances resize --machine-type n2-standard-8 without stopping
AnswerA

Stopping the instance first moves it to the TERMINATED state, where the underlying vCPU/memory allocation can be changed. The `gcloud compute instances set-machine-type` command only works on a stopped instance, so stopping, changing, then starting is the documented, supported path. This preserves the boot disk, persistent disks, static IP, and all instance metadata.

Why this answer

To change the machine type of a running Compute Engine VM, you must stop the instance, use the set-machine-type command (or console) to change the machine type, and then start the instance. This is because the machine type determines the virtual hardware, and changing it requires the instance to be stopped.

Exam trap

ACE often tests the correct procedure for changing machine types, and candidates may think it can be done live or via a resize command, but it requires a stop/start cycle.

How to eliminate wrong answers

Option B is wrong because there is no 'update --machine-type' command for a running instance; the machine type cannot be changed while the instance is running. Option C is wrong because deleting and recreating the instance would lose the instance's configuration and potentially data unless you recreate from a disk, but it's not the correct procedure. Option D is wrong because there is no 'resize' command for machine type; the correct command is 'set-machine-type' and it requires the instance to be stopped.

71
MCQmedium

A Cloud Run service named 'my-service' is currently serving 100% traffic to revision 'rev1'. You deploy a new revision 'rev2' and want to gradually shift traffic so that rev2 receives 10% of requests. Which command should you use?

A.gcloud run services update-traffic my-service --to-revisions=rev2=10,rev1=90
B.gcloud run services update my-service --traffic=rev2=10%
C.gcloud run deploy my-service --image=... --traffic=rev2=10
D.gcloud run revisions update rev2 --traffic=10
AnswerA

This is the correct service-level command for a precise traffic split between two existing revisions. It specifies both revisions explicitly, so Cloud Run routes exactly 10% of requests to rev2 and 90% to rev1; percentages must sum to 100 and should be entered as bare integers (no '%' sign). Because it uses `--to-revisions` on `update-traffic`, it works for rollbacks and gradual shifts without creating a new revision.

Why this answer

`gcloud run services update-traffic` is the dedicated command for changing traffic allocation between revisions, and the `--to-revisions=rev2=10,rev1=90` syntax assigns 10% to rev2 and 90% to rev1. This performs a gradual traffic shift without redeploying the service. It is the correct way to implement a canary or gradual rollout on Cloud Run.

Exam trap

ACE often tests the specific subcommand for traffic management, so candidates who use the generic `gcloud run services update` or `deploy` with a `--traffic` flag pick an invalid syntax.

How to eliminate wrong answers

Option B is wrong because `gcloud run services update` does not accept a `--traffic` flag with percentage syntax; traffic changes require `update-traffic`. Option C is wrong because `gcloud run deploy` with `--traffic` is not the documented way to split traffic between existing revisions and would redeploy rather than adjust routing. Option D is wrong because `gcloud run revisions update` is not a valid command for setting traffic percentages on a service.

72
MCQmedium

You are troubleshooting a Pub/Sub subscription that is not receiving messages as fast as they are published. You want to check if there is a backlog of unacknowledged messages for the subscription. What should you use?

A.Check the Cloud Logging logs for the subscription
B.Use gcloud pubsub subscriptions describe and check the ackDeadlineSeconds
C.Check the Cloud Console Pub/Sub dashboard for the topic publish rate
D.Use Cloud Monitoring to view the 'oldest_unacked_message_age' metric
AnswerD

The `oldest_unacked_message_age` metric, available in Cloud Monitoring under `pubsub.googleapis.com/subscription/oldest_unacked_message_age`, is a gauge metric that reports, per subscription, the age of the oldest message that has not yet been acknowledged. A high or increasing value directly indicates that the subscriber is not keeping up with the message flow, representing a growing backlog. This is the standard and most direct way to detect consumer lag in Cloud Pub/Sub, making it the correct diagnostic tool for the scenario.

Why this answer

The 'oldest_unacked_message_age' metric in Cloud Monitoring directly measures the age of the oldest unacknowledged message in a subscription, which is the definitive indicator of a backlog. A growing value means messages are accumulating faster than they are being acknowledged. This metric is part of the standard Pub/Sub monitoring metrics and is the recommended way to detect and quantify subscription backlog.

Exam trap

ACE often tests the difference between configuration details (like ackDeadlineSeconds) and operational metrics (like oldest_unacked_message_age), so candidates may mistakenly choose a configuration check when asked to diagnose a backlog.

How to eliminate wrong answers

Option A is wrong because Cloud Logging captures operational logs (e.g., publish/ack events, errors) but does not provide a direct metric for backlog size or message age. Option B is wrong because 'gcloud pubsub subscriptions describe' returns configuration details like ackDeadlineSeconds, not real-time backlog data; ackDeadlineSeconds is a configuration setting, not a measure of unacknowledged messages. Option C is wrong because the Cloud Console Pub/Sub dashboard shows topic publish rate, which indicates incoming message volume but not the subscription's ability to keep up or the backlog of unacknowledged messages.

73
MCQmedium

You need to create a log-based metric that counts the number of errors in your application logs. What must you do first in Cloud Logging?

A.Create an alerting policy with a condition
B.Create a log sink that exports logs to BigQuery
C.Define a filter that matches the error logs
D.Install the Logging agent on your VMs
AnswerC

The correct approach is to define a filter expression in Cloud Logging that matches the error logs (e.g., severity=ERROR or specific text), then use that filter to create a logs-based metric. The metric counter increments for every matching log entry, and the filter becomes the metric's definition, allowing you to alert on the count over time.

Why this answer

In Cloud Logging, a log-based metric is based on a filter. You define the filter using the logging query language to match the logs you want to count, then create the metric from that filter.

74
MCQhard

Your Cloud Run service is receiving a sudden spike in traffic. You want to ensure that the number of concurrent requests per container instance does not exceed 10 to avoid overloading the backend. Which configuration should you set?

A.Set --timeout to 10 seconds
B.Set --concurrency to 10
C.Set --max-instances to 10
D.Set --cpu-throttling to true
AnswerB

Correct: --concurrency controls the maximum number of simultaneous requests that each container instance can process at the same time. With a spike, setting it to 10 means an instance will accept only 10 in-flight requests and Cloud Run will automatically spin up additional instances to handle the remaining traffic, preventing any single instance from being overwhelmed. It directly manages per-instance load rather than total capacity or request duration.

Why this answer

In Google Cloud Run, the --concurrency flag controls the maximum number of simultaneous requests that a single container instance can handle. Setting --concurrency to 10 ensures that no container instance processes more than 10 concurrent requests, which directly addresses the requirement to prevent backend overload. This is the precise configuration parameter designed for this purpose.

Exam trap

ACE often tests the confusion between concurrency (requests per instance) and max-instances (total instances), and candidates frequently pick --max-instances when the question asks about per-instance request limits.

How to eliminate wrong answers

Option A is wrong because --timeout controls how long a request can run before being terminated (default 300 seconds), not how many concurrent requests a container handles. Option C is wrong because --max-instances limits the total number of container instances that Cloud Run can scale to, which controls horizontal scaling, not per-instance concurrency. Option D is wrong because --cpu-throttling is not a valid Cloud Run configuration flag; CPU allocation in Cloud Run is managed through --cpu and --memory settings, and throttling behavior is controlled by whether CPU is always allocated or only during request processing.

75
MCQmedium

Your team manages a Compute Engine instance group that runs a stateless web application. You need to ensure that instances are automatically repaired if they fail health checks, and that the group scales based on CPU utilization. Which type of instance group should you use?

A.Regional managed instance group
B.Instance group with autoscaling only
C.Unmanaged instance group
D.Managed instance group
AnswerD

A managed instance group (MIG) automatically creates and manages instances based on an instance template. It supports autohealing by recreating instances that fail health checks, and it can autoscale based on CPU utilization or other metrics. This directly meets the requirement for automatic repair and scaling, making it the correct choice for the stateless web application.

Why this answer

A managed instance group (MIG) is designed to automate the lifecycle of identical instances. It uses an instance template to create instances and can be configured with autohealing based on health checks, as well as autoscaling based on metrics like CPU utilization. This makes it the ideal choice for a stateless web application that requires automatic repair and scaling.

Exam trap

The trap here is assuming that any instance group can autoheal, when in fact only managed instance groups provide that capability.

Page 1 of 2 · 81 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Ace Ensuring Operation questions.