Courseiva

Google Professional Data Engineer (PDE) — Questions 451–525

747 questions total · 10pages · All types, answers revealed

Page 6

Page 7 of 10

Page 8
451
MCQeasy

A data engineer needs to process streaming data from thousands of IoT devices and generate real-time dashboards. The data volume is low but requires exactly-once processing semantics. Which Google Cloud service combination should they use?

A.Cloud Pub/Sub + Cloud Data Fusion
B.Cloud Pub/Sub + Cloud Dataproc
C.Cloud Pub/Sub + Cloud Dataflow
D.Cloud Pub/Sub + Cloud Dataprep
AnswerC

Cloud Pub/Sub ingests the device streams, while Dataflow provides exactly-once processing through its streaming engine's state and deduplication. This satisfies the stem's low-volume, real-time dashboard requirement, since Dataflow's windowing and triggers emit results continuously rather than only on batch completion.

Why this answer

Dataflow supports exactly-once processing via its streaming engine and checkpointing. Pub/Sub is the ingest service. Together they provide the required semantics.

452
Multi-Selectmedium

A data scientist wants to use Vertex AI Workbench for exploratory data analysis. Which TWO statements are true about Vertex AI Workbench?

Select 2 answers
A.It is a serverless service that scales to zero when not in use.
B.It supports custom container images for the notebook environment.
C.It can only be used with TensorFlow.
D.It provides a managed JupyterLab environment with pre-installed ML libraries.
E.It includes a built-in SQL query editor for BigQuery.
AnswersB, D

Custom container images let the notebook instance run with bespoke dependencies, frameworks or system packages beyond the default image. This satisfies reproducibility and version-pinning needs during exploratory data analysis when specific library builds are required.

Why this answer

Option B is correct because Vertex AI Workbench instances let you specify a custom container image for the notebook environment, so you can bring your own dependencies and tooling rather than being limited to the default image. Option D is correct because Vertex AI Workbench provides a managed JupyterLab environment with pre-installed ML libraries (such as TensorFlow, PyTorch, and scikit-learn), which is exactly what a data scientist needs for exploratory data analysis. Option A is incorrect because Workbench instances are user-managed VMs that do not scale to zero; they keep running (and incurring cost) until stopped.

Option C is incorrect because Workbench supports multiple frameworks and languages, not only TensorFlow. Option E is incorrect because a built-in SQL query editor for BigQuery is a feature of BigQuery and other tools, not a defining capability of Vertex AI Workbench.

Exam trap

PDE often tests the distinction between serverless and managed services, so the trap is assuming Workbench scales to zero like Cloud Functions, when it actually runs on persistent VMs.

453
Drag & Dropmedium

Drag and drop the steps to create a Cloud Composer environment for Apache Airflow into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Cloud Composer provides a managed Airflow environment for orchestrating workflows.

454
MCQhard

A team is using BigQuery to analyze petabyte-scale data. They notice that queries are slow and expensive due to full table scans. They have already partitioned by date. What additional optimization should they implement?

A.Use materialized views
B.Cluster by frequently filtered columns
C.Convert to native tables
D.Use query caching
AnswerB

Clustering sorts and co-locates data by the specified columns within each partition, so BigQuery scans only the relevant blocks when those columns are filtered. This reduces bytes read and cost beyond partitioning alone, directly addressing the full-table-scan problem.

Why this answer

Clustering by frequently filtered columns (option B) organizes data within each partition based on the sort order of those columns. This allows BigQuery to prune blocks during query execution, significantly reducing the amount of data scanned and improving both performance and cost. Since the table is already partitioned by date, clustering adds a secondary ordering that targets the most common filter predicates, avoiding full table scans within each partition.

Exam trap

Google Cloud often tests the distinction between partitioning and clustering, where candidates mistakenly believe partitioning alone is sufficient for all filtering scenarios, but clustering is required to avoid full scans on non-date columns.

How to eliminate wrong answers

Option A is wrong because materialized views precompute and store query results, which can speed up repeated aggregations but do not reduce the scan cost of ad-hoc filters on raw data; they are not a substitute for physical data organization like clustering. Option C is wrong because BigQuery tables are already native (managed) tables; converting to native tables is not a valid operation and does not address scan efficiency. Option D is wrong because query caching only returns results for identical queries run within 24 hours, but it does not reduce the scan cost or improve performance for new or slightly different queries that still trigger full table scans.

455
MCQhard

A healthcare company stores patient imaging metadata in Cloud Firestore in Native mode. The application must run a query that returns all documents where the status field equals 'pending' and the region field equals 'us-east', ordered by created_at descending, and it must paginate through potentially millions of matching documents. Which Firestore capability should the engineer use to satisfy this requirement efficiently?

A.A collection group query across all subcollections with an aggregation count.
B.A single-field index on created_at with offset-based pagination using the skip parameter.
C.A composite index on status, region, and created_at, combined with cursor-based pagination.
D.A distributed counter shard to track pending documents per region.
AnswerC

Firestore requires a composite index for queries combining equality filters on multiple fields with an orderBy on a different field. Cursor-based pagination using startAfter with the last document snapshot avoids the cost and inconsistency of offset-based paging, and it scales to millions of matching documents without re-reading skipped results.

Why this answer

Firestore automatically indexes single fields but requires a composite index when a query combines multiple equality filters with an orderBy on another field. Cursor-based pagination via startAfter scales to large result sets because it resumes from the last document rather than skipping offsets, making it the efficient, correct approach for this query pattern.

Exam trap

The trap here is assuming single-field indexes or offset pagination can handle a multi-filter ordered query, when Firestore requires a composite index and cursors for scale.

456
MCQhard

A data pipeline ingests sensor data from IoT devices via Cloud Pub/Sub, processes it with Cloud Dataflow, and writes to BigQuery. The pipeline is failing with high latency and data loss. Which troubleshooting step should be taken first?

A.Check Stackdriver logging for error messages.
B.Disable exactly-once processing in Dataflow.
C.Increase the number of Dataflow workers.
D.Switch to BigQuery streaming inserts.
AnswerA

Checking Stackdriver (Cloud Logging) first surfaces the pipeline's own error messages, which pinpoint whether latency and loss stem from Pub/Sub backlog, Dataflow worker exhaustion, or BigQuery write rejections. This satisfies the stem's "first" constraint by gathering diagnostic evidence before altering any component.

Why this answer

Stackdriver (now Cloud Logging) is the first place to investigate when a Dataflow pipeline experiences high latency and data loss. Dataflow automatically logs errors, worker failures, and system messages to Cloud Logging, which can reveal root causes such as insufficient resources, stuck steps, or Pub/Sub subscription issues. Checking logs first avoids premature scaling or configuration changes that may not address the actual problem.

Exam trap

Google Cloud often tests the principle of 'diagnose before you optimize' — the trap here is that candidates jump to scaling or switching technologies (options C and D) without first checking logs, which is the fundamental first step in any troubleshooting workflow.

How to eliminate wrong answers

Option B is wrong because disabling exactly-once processing in Dataflow would not fix high latency or data loss; it could actually increase data duplication and make debugging harder, while the core issue remains unaddressed. Option C is wrong because increasing the number of Dataflow workers without first diagnosing the bottleneck (e.g., a hot key, slow transform, or Pub/Sub backlog) can waste resources and may not resolve the underlying cause of latency or loss. Option D is wrong because switching to BigQuery streaming inserts does not address pipeline-level failures; streaming inserts have their own quotas, error handling, and latency characteristics, and the problem likely lies in the Dataflow processing logic or resource allocation, not the sink.

457
MCQhard

A Dataflow streaming pipeline processes events from Pub/Sub and writes to BigQuery using a dynamically generated table destination based on the event type. The pipeline is experiencing high latency, and the worker CPU utilization is low. Which action is most likely to reduce latency?

A.Increase the batch size parameter in the BigQuery sink to write larger batches.
B.Reduce the number of workers to increase CPU utilization per worker.
C.Enable Dataflow Streaming Engine to improve throughput and reduce latency.
D.Increase the worker disk size to reduce I/O wait time.
AnswerC

Enabling Dataflow Streaming Engine moves state and computation from worker VMs to the backend service, reducing per-worker overhead and alleviating shuffle/state bottlenecks. This directly addresses the low CPU utilization and high latency, improving throughput.

Why this answer

Dataflow Streaming Engine moves state and computation from worker VMs to the backend service, reducing per-worker overhead and enabling better resource utilization. This directly addresses the symptom of high latency with low CPU utilization, which indicates workers are bottlenecked on shuffle or state management rather than compute.

Exam trap

The trap here is that candidates often assume low CPU utilization means workers are underutilized and should be scaled down (Option B), when in fact low CPU with high latency indicates a bottleneck in shuffle or state management that is not compute-bound.

How to eliminate wrong answers

Option A is wrong because increasing batch size in the BigQuery sink can actually increase latency for streaming pipelines, as larger batches require more time to fill before writing, and the issue here is not sink throughput but worker inefficiency. Option B is wrong because reducing the number of workers would decrease parallelism and likely worsen latency, and low CPU utilization suggests workers are not compute-bound but rather waiting on I/O or shuffle. Option D is wrong because increasing worker disk size does not reduce I/O wait time for streaming pipelines; disk I/O is not the bottleneck when CPU is low and the pipeline uses Pub/Sub and BigQuery, which are network-bound.

458
Multi-Selectmedium

A company is migrating their on-premises data warehouse to BigQuery. They have a mix of batch and streaming ingestion. The data team wants to optimize query costs. Which THREE practices should they adopt?

Select 3 answers
A.Switch to flat-rate pricing to cap slot usage.
B.Use materialized views for frequently executed aggregations.
C.Partition tables by a date or timestamp column.
D.Limit the number of concurrent queries by setting a maximum slot capacity.
E.Cluster tables on columns that are frequently used in filters and joins.
AnswersB, C, E

Materialised views precompute and store aggregation results, so repeated queries scan the small stored result rather than the full base tables. This directly satisfies the cost-optimisation goal by reducing bytes processed for frequently executed aggregations.

Why this answer

Option B is correct because materialized views precompute and store the results of frequently executed aggregations, so repeated queries read the smaller cached result set instead of rescanning the full base tables, directly reducing bytes processed and thus query cost. Option C is correct because partitioning a table by a DATE or TIMESTAMP column enables partition pruning, so queries with filters on that column scan only the relevant partitions rather than the entire table, cutting the data billed per query. Option E is correct because clustering on columns frequently used in filters and joins co-locates related data in storage blocks, allowing BigQuery to skip irrelevant blocks and reduce bytes scanned for those predicates.

Option A is not correct because flat-rate pricing changes how slots are billed (a committed capacity model) rather than reducing the bytes processed by queries, so it does not itself optimize query cost for this migration. Option D is not correct because capping concurrent queries or slot capacity is a workload-management throttle, not a cost-optimization practice, and it can even slow throughput without lowering bytes-scanned charges.

Exam trap

PDE often tests the confusion between cost optimization and performance/capacity management; candidates pick flat-rate or concurrency limits thinking they reduce cost, but only reducing bytes scanned via partitioning, clustering, and materialized views lowers on-demand query costs.

459
MCQeasy

You need to react to changes in a GCS bucket (e.g., new object creation) and trigger a Cloud Run service to process the new file. Which Google Cloud service should you use to route the event?

A.Pub/Sub directly with a Cloud Run subscription
B.Cloud Tasks
C.Eventarc
D.Cloud Scheduler
AnswerC

Eventarc routes Cloud Storage object-creation events to Cloud Run, satisfying the requirement to trigger the service on new files. It delivers CloudEvents from Google Cloud sources directly to serverless endpoints, unlike Pub/Sub alone, which requires custom wiring, or Cloud Scheduler, which polls rather than reacting to bucket changes.

Why this answer

Eventarc is the correct choice because it is purpose-built to route events from Google Cloud sources (like Cloud Storage) to Cloud Run. It directly supports Cloud Storage audit logs and Pub/Sub event triggers, allowing you to react to object creation events without custom middleware. Eventarc handles the event routing, filtering, and delivery to your Cloud Run service automatically.

Exam trap

Google often tests the misconception that Pub/Sub is the direct answer for any event routing, but the trap here is that Eventarc is the managed service that simplifies the integration between GCS and Cloud Run, making it the correct choice over raw Pub/Sub.

How to eliminate wrong answers

Option A is wrong because Pub/Sub directly with a Cloud Run subscription requires you to manually configure a Pub/Sub topic and subscription, and Cloud Run can only pull messages via a push subscription; Eventarc abstracts this complexity and provides native integration with Cloud Storage events. Option B is wrong because Cloud Tasks is a task queue for asynchronous execution of HTTP requests, not designed for event-driven routing from GCS; it would require you to manually publish tasks in response to events, adding unnecessary overhead. Option D is wrong because Cloud Scheduler is a cron job scheduler for periodic tasks, not an event router; it cannot react to real-time object creation events in a GCS bucket.

460
Matchingmedium

Match each machine learning term to its description.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Model trained on labeled data

Model trained on unlabeled data

Agent learns by interacting with environment

Model performs well on training data but poorly on new data

Why these pairings

In machine learning, supervised learning uses labeled data (inputs mapped to known outputs), unsupervised learning finds patterns in unlabeled data, features are input variables, and labels are the target outputs. Common confusions include mixing supervised/unsupervised definitions or swapping features and labels.

461
Multi-Selecthard

Which THREE best practices should be followed when designing a Dataflow pipeline for real-time data processing?

Select 3 answers
A.Set up monitoring alerts for system lag and data freshness.
B.Use static side inputs that are loaded once at pipeline start.
C.Implement watermark estimation to handle late data.
D.Use global windows with early triggers for low latency.
E.Use idempotent sinks to ensure exactly-once processing.
AnswersA, C, E

Real-time pipelines accumulate lag between event ingestion and output, so alerting on system lag and data freshness surfaces backpressure or stalled processing before downstream consumers act on stale results. This satisfies the stem's real-time processing requirement by detecting latency regressions continuously rather than after batch completion.

Why this answer

Option A is correct because monitoring system lag (the difference between event time and processing time) and data freshness is essential for detecting pipeline backlog and ensuring timely real-time results in Dataflow. Option C is correct because watermark estimation is how Dataflow tracks event-time completeness and determines when to fire window results while gracefully handling late-arriving data. Option E is correct because idempotent sinks allow retries and replays to produce the same result, which is required to achieve exactly-once processing semantics in streaming pipelines.

Option B is not appropriate because static side inputs loaded once at pipeline start cannot reflect changing reference data in a real-time stream, where dynamic or periodically refreshed side inputs are needed. Option D is not a best practice because global windows with early triggers discard windowing structure and can produce incorrect or incomplete aggregations for unbounded data, whereas proper event-time windowing with allowed lateness is preferred.

Exam trap

Google Cloud often tests the misconception that static side inputs are acceptable for streaming pipelines, but they are only appropriate for batch or bounded data; real-time pipelines require side inputs that can be periodically refreshed (e.g., via a streaming source or a periodic lookup).

462
Multi-Selecthard

A data analyst wants to compute the rank of sales per region and also the difference in sales between consecutive months for each region. Which BigQuery analytic functions should they use? (Select TWO)

Select 2 answers
A.RANK()
B.ROW_NUMBER()
C.LAG()
D.LEAD()
E.NTILE()
AnswersA, C

RANK() is a navigation function that assigns a rank to each row within a partition, ordered by a specified expression. Partitioning by region and ordering by sales yields per-region sales rankings, satisfying the ranking half of the requirement; a separate function handles consecutive-month differences.

Why this answer

RANK() (option A) is correct because it assigns a rank to each row within a partition ordered by sales, with tied sales values receiving the same rank and subsequent ranks skipped, which directly satisfies the requirement to compute the rank of sales per region. LAG() (option C) is correct because it accesses the value of sales from the preceding row within the same region partition ordered by month, allowing the analyst to compute the difference in sales between consecutive months. ROW_NUMBER() (option B) is not appropriate here because it assigns unique sequential numbers even for tied sales values, which does not produce a true rank.

LEAD() (option D) looks forward to the next row rather than backward, so it would not give the previous month's sales for a consecutive-month difference. NTILE() (option E) divides rows into a specified number of buckets, which is unrelated to ranking sales or computing month-over-month differences.

463
Multi-Selectmedium

You want to optimize BigQuery costs for a large dataset that is frequently queried by time range. You also need to ensure that predictable workloads have dedicated slot capacity. Which TWO strategies should you combine? (Choose 2)

Select 2 answers
A.Use query caching
B.Partition the table by date
C.Purchase committed use reservations for baseline capacity
D.Create a materialized view for the entire table
E.Enable autoscaling slots
AnswersB, C

Partitioning by date restricts each query to the relevant date range, so BigQuery scans only matching partitions rather than the whole table. This directly cuts bytes processed for time-range queries, satisfying the cost-optimisation requirement in the stem.

Why this answer

Option B is correct because partitioning the table by date lets BigQuery prune partitions and scan only the date ranges a query touches, which directly reduces bytes processed and therefore cost for a dataset that is frequently queried by time range. Option C is correct because committed use reservations (capacity commitments) provide dedicated, predictable slot capacity at a discounted rate for steady baseline workloads, satisfying the requirement that predictable workloads have dedicated slots. Option A is not appropriate as a primary cost strategy here because query caching only helps when identical queries are repeated and results are still cached, which does not address time-range pruning or dedicated capacity.

Option D is wrong because a materialized view over the entire table would be expensive to maintain and does not target the time-range access pattern. Option E is wrong because autoscaling slots add flexible, on-demand capacity rather than the dedicated, predictable capacity the scenario requires.

Exam trap

The trap is picking autoscaling slots for 'predictable workloads' — autoscaling is for variable demand, while committed use reservations are the correct answer for dedicated, predictable baseline capacity.

464
Multi-Selecthard

Your company runs a Dataflow streaming pipeline that processes user activity from Pub/Sub and writes aggregated results to BigQuery. Lately, the pipeline is experiencing high latency and backlog growth during peak hours. You need to troubleshoot and improve performance. Which THREE actions should you take? (Choose 3.)

Select 3 answers
A.Change the worker machine type to a higher CPU/memory configuration
B.Decrease the window duration to reduce data per window
C.Enable Dataflow Streaming Engine
D.Increase the number of workers in the pipeline
E.Add additional Pub/Sub subscriptions to the same topic
AnswersA, C, D

Scaling worker machine type raises per-worker CPU and memory, directly relieving the compute saturation causing backlog growth at peak. Dataflow autoscaling adds workers, but each remains bound by its machine's vCPU and memory limits, so upgrading the type increases throughput per worker for CPU-bound aggregation stages.

Why this answer

Option A is correct because moving to a higher CPU/memory worker machine type gives each worker more compute and memory to process elements, which directly addresses CPU or memory saturation that causes high latency and backlog growth during peak hours. Option C is correct because enabling Dataflow Streaming Engine moves pipeline state and windowing execution off the worker VMs to the Dataflow service, reducing worker CPU/memory pressure and improving streaming performance and autoscaling responsiveness. Option D is correct because increasing the number of workers adds horizontal processing capacity, allowing the pipeline to drain the Pub/Sub backlog faster and keep up with peak throughput.

Option B is not appropriate because shortening the window duration does not reduce the total volume of data processed and can actually increase per-window overhead and output frequency, potentially worsening latency. Option E is not appropriate because adding more subscriptions to the same topic does not increase the throughput of the existing pipeline; each subscription would need its own reader/pipeline, and extra subscriptions only add fan-out consumers rather than improving the current pipeline's performance.

Exam trap

The trap is assuming that more Pub/Sub subscriptions or smaller windows improve throughput — both actually increase overhead or duplicate data rather than scaling the pipeline.

465
MCQhard

A healthcare startup is deploying a natural language processing (NLP) model for extracting medical entities from clinical notes. The model is a fine-tuned BERT model served on Vertex AI Prediction using a custom container. The team observes that prediction latency is around 500ms per request, but they need to handle up to 100 requests per second (QPS) with end-to-end latency under 200ms. The model currently runs on n1-standard-4 machines (4 vCPU, 15 GB memory). During load testing, CPU utilization reaches 90% and memory usage is 12 GB. The team is considering options to meet the requirements. Which action should they take?

A.Use a machine type with a GPU, such as n1-standard-4 with a NVIDIA Tesla T4 accelerator, and optimize the model with TensorRT.
B.Switch to n1-highmem-4 machines to provide more memory for the model.
C.Deploy the model using TensorFlow Serving with CPU-only nodes and increase the number of replicas.
D.Move the model to Cloud Run with automatic scaling to handle the QPS.
AnswerA

CPU is saturated at 90% with 12 GB memory used, so horizontal scaling alone cannot cut per-request latency. A T4 GPU offloads BERT's matrix multiplications, and TensorRT fuses layers and applies FP16 kernels, cutting inference time well below the 200ms end-to-end target at 100 QPS.

Why this answer

The bottleneck is CPU-bound inference (90% CPU utilization) with memory well within limits (12 GB of 15 GB). Adding a GPU (NVIDIA Tesla T4) and optimizing with TensorRT reduces per-request latency via hardware acceleration and graph optimizations, enabling sub-200ms inference at 100 QPS. This directly addresses the latency requirement without changing the machine family or scaling strategy.

Exam trap

Google Cloud often tests the misconception that scaling horizontally (more replicas or Cloud Run) solves latency problems, when the real issue is per-request compute bottleneck that requires hardware acceleration or model optimization.

How to eliminate wrong answers

Option B is wrong because memory is not the bottleneck (12 GB used out of 15 GB); increasing memory does not reduce CPU-bound inference latency. Option C is wrong because TensorFlow Serving on CPU-only nodes still relies on CPU compute, and increasing replicas adds cost and complexity without addressing the fundamental latency per request; the CPU utilization is already saturated, so more replicas would require horizontal scaling but still not guarantee sub-200ms latency per request. Option D is wrong because Cloud Run's automatic scaling handles QPS but does not reduce per-request latency; the model's inference time remains CPU-bound, and Cloud Run's cold starts and CPU-only instances would not meet the 200ms latency target.

466
MCQmedium

A company wants to use dbt to transform data in BigQuery. Their source data is loaded daily into staging tables. They need to run dbt transformations on a schedule and only process tables that have changed. Which dbt feature should they use?

A.dbt snapshots
B.dbt incremental models
C.dbt seeds
D.dbt sources
AnswerB

dbt incremental models process only new or changed rows since the last run, using a configured unique key and filter. This satisfies the requirement to run on a schedule while avoiding full reprocessing of unchanged staging tables.

Why this answer

dbt incremental models are designed to process only new or changed data since the last run, rather than rebuilding the entire table. This is ideal for daily transformations on staging tables where only some data changes. Incremental models use a unique key and a strategy (e.g., merge, delete+insert) to update the target table efficiently.

Exam trap

The trap is confusing incremental models with snapshots. Candidates might think snapshots process only changed data, but snapshots are for tracking history, not for efficient incremental updates.

How to eliminate wrong answers

Option A is wrong because dbt snapshots are used to track historical changes (SCD Type 2) by capturing changes over time, not for incremental processing of source data. Option C is wrong because dbt seeds are CSV files loaded into the data warehouse, typically for static reference data, not for transforming large datasets. Option D is wrong because dbt sources are configurations that define and document source tables, but they do not provide incremental processing logic.

467
Multi-Selectmedium

Your company uses Cloud Composer to orchestrate a data pipeline that includes Dataproc Spark jobs and BigQuery load operations. You need to pass the output file path from the Spark job to the next BigQuery task in the DAG. Which two mechanisms can you use to share data between tasks? (Choose TWO.)

Select 2 answers
A.Store the output path as a Cloud Composer variable.
B.Publish the output path to a Pub/Sub topic and subscribe in the next task.
C.Write the output path to a Cloud Storage object and read it in the next task.
D.Use BigQuery as an intermediary to store the output path.
E.Use Airflow XComs to push the output path from the Spark task and pull it in the BigQuery task.
AnswersC, E

Writing the output path to a Cloud Storage object and reading it in the next task decouples the Dataproc Spark job from the BigQuery load, persisting the value outside the DAG run. This satisfies the requirement to pass the file path reliably between tasks.

Why this answer

Option E is correct because Airflow XComs are the native, built-in mechanism for passing small pieces of data such as a file path between tasks in a DAG: the Spark task calls xcom_push (or returns a value) and the BigQuery task uses xcom_pull with the task_id/key to retrieve it. Option C is correct because writing the output path to a Cloud Storage object and reading it in the next task is a durable, decoupled way to share data between tasks, especially when the value is larger or must persist beyond the DAG run; Cloud Composer tasks can use the GCS operators/sensors or the google-cloud-storage client to write and read the object. Option A is not appropriate because Cloud Composer (Airflow) Variables are global key-value configuration stores, not per-run task-to-task data channels, so using them for a run-specific output path risks race conditions and stale values.

Option B is not appropriate because Pub/Sub is an asynchronous messaging service, not a synchronous task-to-task handoff mechanism, and it adds unnecessary complexity and latency for passing a single path within one DAG run. Option D is not appropriate because using BigQuery as an intermediary to store a transient output path is heavyweight and semantically wrong; BigQuery is for analytical data, not for orchestrating task metadata.

Exam trap

PDE often tests the misconception that global variables or external services like Pub/Sub are appropriate for task-to-task communication, when in fact Airflow XComs and Cloud Storage are the canonical mechanisms.

468
MCQmedium

Your organization runs a Cloud Composer (Apache Airflow) environment that executes a DAG every night to move data from Cloud Storage into BigQuery. The DAG has been succeeding for months, but last week the nightly load silently produced a BigQuery table with zero rows while the Airflow task still reported success. You need to add a safeguard that fails the DAG task whenever the loaded row count is zero before downstream tasks run. What should you do?

A.Increase the retries parameter on the load task so transient issues that produce zero rows are automatically retried.
B.Add a BigQueryCheckOperator task immediately after the load task that asserts the target table row count is greater than zero.
C.Enable Cloud Logging export for the Composer environment and create a log-based alert that matches error text from the load operator.
D.Configure an Airflow SLA on the load task so that a missed deadline raises an alert about the empty table.
AnswerB

The BigQueryCheckOperator runs a SQL statement and fails the task when the returned result does not match the expected condition. Placing it right after the load task means a zero-row result raises an Airflow task failure, stopping downstream tasks and surfacing the problem through the normal DAG alerting path instead of silently completing.

Why this answer

The failure mode is a task that reports success while producing an empty result set, so the safeguard must inspect the data itself and raise a task-level failure. A BigQueryCheckOperator performs exactly that assertion and integrates with Airflow's dependency graph, which prevents downstream tasks from consuming invalid output and triggers configured alerting.

Exam trap

The trap here is assuming that any task marked successful in Airflow implies the underlying data operation was semantically correct, when success only reflects that the operator did not raise an exception.

469
MCQmedium

Refer to the exhibit. A team configured a Cloud Monitoring alerting policy as shown. They recently started receiving false positive alerts. What is the most likely cause?

A.The duration of 60 seconds is too short, making the alert sensitive to brief spikes.
B.The alignment period of 60 seconds is too short, causing noise.
C.The threshold of 10 is too low.
D.The aggregator should be ALIGN_SUM instead of ALIGN_RATE.
AnswerA

A 60-second duration means the condition need only hold briefly before firing, so short transient spikes breach the threshold and trigger alerts. Lengthening duration requires sustained violation, filtering brief spikes that do not represent genuine incidents.

Why this answer

A 60-second duration means the alert fires if the condition is met for just one minute. This is too short to distinguish transient spikes from sustained issues, causing false positives. Increasing the duration would require the metric to breach the threshold for a longer, more meaningful period.

Exam trap

Google often tests the distinction between alignment period (how data is aggregated) and duration (how long the condition must persist), tempting candidates to blame the alignment period when the real issue is the insufficient duration.

How to eliminate wrong answers

Option B is wrong because the alignment period of 60 seconds is standard for aggregating data into regular intervals; a shorter alignment period could increase noise, but the primary cause of false positives here is the short duration, not the alignment. Option C is wrong because a threshold of 10 is not inherently too low; the false positives are due to the alert triggering on brief spikes, not because the threshold value is misconfigured. Option D is wrong because ALIGN_RATE is appropriate for metrics that measure change over time (e.g., requests per second), and using ALIGN_SUM would sum rates incorrectly, potentially masking spikes rather than causing false positives.

470
MCQeasy

A company wants to stream real-time clickstream data from a website into BigQuery for near-real-time analytics. They expect peaks of 10,000 events per second. Which combination of services is most suitable for ingestion?

A.Cloud Storage → Cloud Functions → BigQuery
B.Direct Web → Dataflow → BigQuery
C.Pub/Sub → Dataflow → BigQuery (Storage Write API)
D.Pub/Sub → Dataflow → BigQuery (legacy streaming inserts)
AnswerC

Pub/Sub decouples ingestion from processing, absorbing 10,000 events per second spikes without loss, while Dataflow provides exactly-once streaming transformation. The Storage Write API commits rows directly into BigQuery, satisfying the near-real-time analytics requirement without staging files in Cloud Storage.

Why this answer

For high-throughput, real-time streaming into BigQuery, the combination of Pub/Sub for ingestion, Dataflow for stream processing, and BigQuery's Storage Write API for writing is the most suitable. Pub/Sub handles the 10,000 events per second peaks reliably, Dataflow provides scalable stream processing, and the Storage Write API offers exactly-once semantics and better performance than legacy streaming inserts.

Exam trap

The trap is assuming that any streaming combination works equally well, but the exam expects knowledge that the Storage Write API is preferred over legacy streaming inserts for performance and exactly-once semantics, and that Pub/Sub is essential for handling peak loads.

How to eliminate wrong answers

Option A is wrong because Cloud Storage → Cloud Functions → BigQuery is a batch-oriented or event-driven approach that does not handle high-throughput real-time streaming well; Cloud Functions have concurrency limits and are not designed for sustained 10k EPS. Option B is wrong because Direct Web → Dataflow → BigQuery bypasses Pub/Sub, losing the buffering and decoupling benefits needed for peak loads; Dataflow alone may struggle with ingestion spikes without a message queue. Option D is wrong because Pub/Sub → Dataflow → BigQuery (legacy streaming inserts) uses the older streaming API, which is more expensive, has lower throughput limits, and does not provide exactly-once semantics, making it less suitable than the Storage Write API.

471
MCQeasy

Which Google Cloud service is a serverless, highly scalable data warehouse for analytical queries, supporting SQL and integration with BI tools?

A.Firestore
B.Cloud SQL
C.Cloud Spanner
D.BigQuery
AnswerD

BigQuery is Google Cloud's fully managed, serverless analytics warehouse, separating storage from compute so it scales automatically. It natively runs ANSI SQL and connects to BI tools such as Looker and Data Studio, directly satisfying the serverless, highly scalable analytical query requirement in the stem.

Why this answer

BigQuery is a serverless, highly scalable data warehouse designed for analytical queries over large datasets. It supports standard SQL and integrates seamlessly with BI tools like Looker and Tableau, making it the correct choice for this use case.

Exam trap

The trap here is that candidates may confuse Google Cloud Spanner's global scale and SQL support with data warehousing capabilities, but Spanner is optimized for transactional consistency, not analytical query performance or BI tool integration.

How to eliminate wrong answers

Option A is wrong because Firestore is a NoSQL document database for mobile and web app development, not a data warehouse for analytical SQL queries. Option B is wrong because Cloud SQL is a fully managed relational database for OLTP workloads, not a serverless data warehouse optimized for large-scale analytics. Option C is wrong because Cloud Spanner is a globally distributed, strongly consistent relational database service for transactional workloads, not a data warehouse designed for analytical queries and BI integration.

472
MCQhard

A data engineer needs to alert when Pub/Sub subscription has messages older than 1 hour. Which Cloud Monitoring metric and filter should they use?

A.Metric: topic/send_message_operation_count; filter: topic_id
B.Metric: subscription/ack_message_count; filter: subscription_id
C.Metric: subscription/num_undelivered_messages; filter: subscription_id
D.Metric: subscription/oldest_unacked_message_age; filter: subscription_id
AnswerD

The `subscription/oldest_unacked_message_age` metric reports the age of the oldest unacknowledged message, directly satisfying the one-hour staleness threshold. Filtering by `subscription_id` scopes the alert to the specific Pub/Sub subscription, so Cloud Monitoring evaluates only that subscription's backlog age rather than aggregating across the project.

Why this answer

The metric subscription/oldest_unacked_message_age directly reports the age of the oldest unacknowledged message in a subscription, which is exactly what's needed to alert when messages are older than 1 hour. The filter subscription_id scopes the metric to the specific subscription. This metric is designed for monitoring message backlog age and is the correct choice for this requirement.

Exam trap

PDE often tests the distinction between message count and message age metrics, tempting candidates to choose num_undelivered_messages when the requirement is about age.

How to eliminate wrong answers

Option A is wrong because topic/send_message_operation_count measures the number of publish operations to a topic, not message age in a subscription. Option B is wrong because subscription/ack_message_count counts acknowledged messages, not the age of unacknowledged ones. Option C is wrong because subscription/num_undelivered_messages gives the count of undelivered messages, not their age, so it cannot detect messages older than 1 hour.

473
Drag & Dropmedium

Drag and drop the steps to set up a Pub/Sub topic with a push subscription to an HTTPS endpoint into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order to set up a Pub/Sub topic with a push subscription to an HTTPS endpoint is to first create the topic, then create a push subscription specifying the HTTPS endpoint URL. This ensures the topic exists before the subscription references it, and the push endpoint is configured at subscription creation time. Common mistakes include creating the subscription before the topic, omitting the endpoint, or attempting to convert a pull subscription to push, which is not supported.

474
Multi-Selecthard

You are designing a streaming pipeline that must guarantee exactly-once processing. Which three services or features can help achieve this? (Choose THREE.)

Select 3 answers
A.Cloud Functions for post-processing
B.BigQuery streaming inserts with a unique key for deduplication
C.Cloud Spanner for deduplication state across the pipeline
D.Cloud Pub/Sub with duplicate detection (using message IDs)
E.Dataflow with idempotent write operations to BigQuery
AnswersC, D, E

Using Cloud Spanner as a global state store allows tracking processed event IDs for deduplication.

Why this answer

Cloud Spanner is correct because it provides globally distributed, strongly consistent transactions that can be used to maintain deduplication state across the entire streaming pipeline. By storing a unique key for each processed event in Spanner, the pipeline can atomically check and record whether an event has already been handled, ensuring exactly-once semantics even in the face of retries or failures.

Exam trap

Google Cloud often tests the misconception that BigQuery streaming inserts can guarantee exactly-once processing via a unique key, when in fact BigQuery only supports at-least-once delivery and requires external deduplication mechanisms like Cloud Spanner or Dataflow with idempotent writes.

475
Drag & Dropmedium

Drag and drop the steps to set up Cloud IAP (Identity-Aware Proxy) for an App Engine app into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

IAP verifies identity and authorization before allowing access to the application.

476
MCQeasy

A company processes CSV files that are uploaded to Cloud Storage by external partners. Each file is around 500 MB, and they need to be parsed and loaded into BigQuery. The processing must start as soon as the file arrives. What is the most efficient serverless architecture?

A.Cloud Storage triggers a Cloud Function that publishes events to Pub/Sub; a Dataflow streaming pipeline reads from Pub/Sub and writes to BigQuery.
B.Use Cloud Scheduler to periodically check for new files and process them with Dataflow batch jobs.
C.Cloud Storage triggers a Dataproc job that reads the file and loads it into BigQuery.
D.Cloud Storage triggers a Cloud Function that directly loads the data into BigQuery using the BigQuery API.
AnswerA

Cloud Storage object-finalise events trigger the Cloud Function, which publishes to Pub/Sub so Dataflow's streaming pipeline begins parsing immediately on arrival. This serverless chain handles 500 MB files without provisioning servers and satisfies the start-on-arrival requirement.

Why this answer

It combines Cloud Storage event-driven triggers with Pub/Sub for reliable asynchronous message delivery, and uses Dataflow streaming with autoscaling to handle 500 MB files efficiently. This serverless architecture ensures processing starts immediately upon file arrival, scales to handle large files without manual intervention, and leverages BigQuery's streaming inserts for near-real-time data loading.

Exam trap

Google Cloud often tests the misconception that Cloud Functions can handle large file processing directly, but the 9-minute timeout and memory limits make them unsuitable for files over a few hundred MB, pushing candidates toward the seemingly simpler Option D.

How to eliminate wrong answers

Option B is wrong because Cloud Scheduler polling introduces latency and inefficiency, as it checks for new files on a fixed schedule rather than reacting instantly, which violates the requirement that processing must start as soon as the file arrives. Option C is wrong because Dataproc is a managed Hadoop/Spark service that requires cluster provisioning and startup time, adding overhead for a simple CSV-to-BigQuery load; it is not serverless and not the most efficient for this use case. Option D is wrong because Cloud Functions have a 9-minute timeout and 2 GB memory limit, making them unsuitable for parsing and loading a 500 MB CSV file directly via the BigQuery API, which would likely exceed these constraints and cause failures.

477
Multi-Selecteasy

A company needs to store and analyze large amounts of unstructured data (images, videos) and structured data (CSV logs) in a cost-effective manner. The data should be accessible for analytics with BigQuery. Which two services should they use? (Choose TWO.)

Select 2 answers
A.Cloud SQL
B.Cloud Spanner
C.BigQuery
D.Cloud Storage
E.Firestore
AnswersC, D

BigQuery provides the serverless analytics engine that queries the data and supports federated or external tables, satisfying the requirement to analyse both structured CSV logs and unstructured files without provisioning infrastructure. It delivers the analytics layer the scenario demands.

Why this answer

Cloud Storage is the best option for storing unstructured and structured files cost-effectively. BigQuery can analyze this data directly via external tables or after loading, making it a powerful analytics platform.

478
MCQmedium

You are monitoring a Dataproc cluster and notice that the cluster utilisation is high, but jobs are running slowly. The cluster uses preemptible workers for cost savings. What is the most likely cause of the performance degradation?

A.The primary workers are using standard disks instead of SSDs.
B.The preemptible workers are being preempted frequently, causing task retries and slowdowns.
C.The cluster is under-provisioned; increase the number of preemptible workers.
D.The cluster is using an older image version; upgrade to the latest.
AnswerB

Frequent preemption of preemptible VMs forces Dataproc to rerun interrupted tasks on remaining workers, inflating job duration despite high CPU utilisation. Since the stem specifies preemptible workers used for cost savings, this directly explains the slowdown: capacity vanishes mid-job, triggering retries and stragglers rather than a genuine resource shortage.

Why this answer

Preemptible (spot) workers in Dataproc can be reclaimed by Compute Engine at any time, and when they are preempted mid-task, the task must be retried on remaining workers. Frequent preemption causes repeated task retries, stragglers, and overall job slowdown even though the cluster appears highly utilized. This is the classic symptom of relying heavily on preemptible workers.

Exam trap

PDE often tests the misconception that high utilization means the cluster is healthy, when in fact preemptible worker churn causes retries that inflate utilization while slowing jobs.

How to eliminate wrong answers

Option A is wrong because standard disks on primary workers would slow I/O but would not produce the intermittent, retry-driven slowdown pattern associated with preemption; also, primary workers are not preempted. Option C is wrong because adding more preemptible workers does not fix the root cause — more preemptible capacity can actually increase the preemption rate and retry churn. Option D is wrong because an older image version is a static configuration issue that would affect all jobs consistently, not cause the fluctuating slowdowns tied to preemption.

479
MCQeasy

A data engineering team maintains several Cloud Composer DAGs that load files from Cloud Storage into BigQuery. They want the DAG to start only after a new file has landed in the source bucket, rather than polling on a fixed schedule and often finding nothing. Which Airflow construct should they use to trigger the DAG based on the arrival of the object?

A.A GCSObjectExistenceSensor with a poke interval, followed by the load task.
B.A TimeDeltaSensor set to the expected upload window before the load task.
C.A BigQueryInsertJobOperator configured to load the file directly from the bucket path.
D.An ExternalTaskSensor pointed at a separate file-upload DAG.
AnswerA

The GCSObjectExistenceSensor waits until the specified object exists in Cloud Storage before allowing the DAG to proceed. This turns a time-driven schedule into an event-aware one: the load task runs only when the file is actually present, eliminating wasted polling runs that find nothing and reducing BigQuery load attempts against empty inputs.

Why this answer

The requirement is to begin processing only when a specific object appears, which is precisely what a Cloud Storage sensor provides. Sensors defer downstream tasks until their condition is met, so pairing a GCSObjectExistenceSensor with the load task converts the DAG from a blind schedule into a data-arrival trigger while keeping the existing load logic intact.

Exam trap

The trap here is confusing a time-based wait with an event-based wait, since both keep a task pending, but only a sensor that inspects Cloud Storage can confirm the object actually exists.

480
MCQmedium

A gaming company uses Avro schemas for its streaming event data. They anticipate adding new optional fields to events over time. They need to ensure backward compatibility so that existing pipelines continue to work. Which strategy should they adopt?

A.Use Avro with a schema registry that enforces backward-compatible changes
B.Use JSON instead of Avro and ignore unknown fields
C.Use Protocol Buffers with breaking changes
D.Use FlatBuffers for performance
AnswerA

Avro's schema evolution rules allow adding optional fields without breaking existing consumers, and a schema registry enables version management.

Why this answer

Avro, combined with a schema registry, allows schema evolution with backward compatibility. The registry enforces rules such as adding optional fields with defaults, ensuring that consumers using older schemas can still deserialize new data without breaking. This directly addresses the requirement for existing pipelines to continue working as new optional fields are added.

Exam trap

Google Cloud often tests the misconception that any serialization format (like JSON or Protocol Buffers) inherently supports backward compatibility, but the key is the combination of a schema registry with enforced evolution rules, which only Avro explicitly provides in this context.

How to eliminate wrong answers

Option B is wrong because JSON lacks a schema enforcement mechanism; while ignoring unknown fields is possible, JSON does not provide built-in compatibility guarantees or schema evolution rules, making it error-prone in large-scale streaming systems. Option C is wrong because Protocol Buffers can support backward compatibility, but the option specifies 'breaking changes,' which would violate the requirement for backward compatibility. Option D is wrong because FlatBuffers prioritize performance (zero-copy deserialization) but do not inherently enforce backward-compatible schema evolution, and they are less suited for streaming event data with frequent schema changes.

481
Multi-Selecthard

Which THREE steps are essential for implementing a continuous training pipeline with Vertex AI?

Select 3 answers
A.If the new model passes evaluation, deploy it to a production endpoint.
B.Manually approve each new model version before deployment.
C.Deploy the original model once and set it to auto-update.
D.Set up a trigger to start a training pipeline when new training data is available (e.g., via Cloud Storage events).
E.Include a step in the pipeline that evaluates the new model against a validation set.
AnswersA, D, E

Passing evaluation gates the promotion, so only a model meeting the defined metric threshold reaches the production endpoint. This conditional deployment step completes the continuous training pipeline by automating release, satisfying the requirement that retrained models serve live traffic without manual intervention.

Why this answer

Option A is correct because a continuous training pipeline must close the loop by promoting a model that passes evaluation to a production endpoint, typically via Vertex AI Model Registry and an Endpoint deployment, so the retrained model actually serves predictions. Option D is correct because continuous training requires an automated trigger, such as a Cloud Storage object-finalize event routed through Eventarc or Cloud Functions to launch the Vertex AI Pipeline when new training data lands. Option E is correct because the pipeline must include a model evaluation step that scores the new model against a held-out validation set (for example using Vertex AI Model Evaluation metrics) to gate promotion.

Option B is not essential because manual approval breaks the continuous/automated nature of the pipeline, even if it can be added as an optional gate. Option C is incorrect because deploying the original model with auto-update is not a Vertex AI mechanism for retraining; continuous training is driven by pipelines and triggers, not by an endpoint auto-updating itself.

Exam trap

Candidates often mistakenly include manual approval (B) as essential in a continuous training pipeline, or believe models can auto-update (C) without explicit pipeline steps. For Vertex AI, the required steps are triggering via events, evaluation, and automated deployment upon passing checks.

482
MCQhard

A data pipeline ingests real-time events from Cloud Pub/Sub into BigQuery using Dataflow. The pipeline uses a sliding window of 5 minutes with a 1-minute period to aggregate event counts. Recently, the pipeline started failing with 'The worker failed to provide a heartbeat.' The Dataflow logs show high CPU usage on the workers. What is the best course of action to resolve the issue?

A.Increase the number of workers and enable autoscaling to distribute the load.
B.Reduce the number of workers to minimize coordination overhead.
C.Use a global window with a trigger to reduce state size.
D.Change the windowing to a fixed 5-minute window to reduce computations.
AnswerA

Scaling out workers with autoscaling directly relieves the CPU saturation causing missed heartbeats, since each worker's load shrinks as the sliding window's per-key aggregation is redistributed. This satisfies the stem's high-CPU constraint without altering window semantics, unlike tuning the 5-minute window or 1-minute period.

Why this answer

The 'worker failed to provide a heartbeat' error combined with high CPU usage indicates that workers are overloaded and cannot process data fast enough to maintain their heartbeat to the Dataflow service. Increasing the number of workers and enabling autoscaling distributes the computational load across more machines, reducing per-worker CPU pressure and allowing heartbeats to be sent on time. This directly addresses the root cause of resource exhaustion.

Exam trap

Google Cloud often tests the misconception that reducing workers or changing window types is a universal fix for resource exhaustion, when in fact the immediate solution for heartbeat failures due to high CPU is to scale out the worker pool.

How to eliminate wrong answers

Option B is wrong because reducing the number of workers would concentrate the same workload on fewer machines, increasing per-worker CPU usage and worsening the heartbeat failure. Option C is wrong because using a global window with a trigger does not reduce state size for sliding windows; it would accumulate all events into a single unbounded window, potentially increasing memory pressure and CPU overhead. Option D is wrong because changing to a fixed 5-minute window does not reduce computations compared to a sliding window with a 1-minute period; it actually changes the semantics (non-overlapping windows) and may still cause high CPU if the underlying load is unchanged.

483
MCQeasy

A data engineer needs to run a recurring nightly extract-transform-load job that pulls data from a REST API, applies Python transformations, and writes the output to a Cloud Storage bucket. The team wants a fully managed, serverless scheduler that can retry failed runs and send notifications, and they do not want to maintain any cluster or VM. Which Google Cloud service should they use to define and run this job?

A.A Compute Engine instance running a cron job that executes a Python script and writes to Cloud Storage.
B.Cloud Composer with a single-node environment and a KubernetesPodOperator that runs the Python transformation.
C.Cloud Scheduler triggering a Cloud Function that performs the API call, transformation, and Cloud Storage write.
D.Dataproc Serverless with a PySpark batch that calls the REST API and writes results to Cloud Storage.
AnswerC

Cloud Scheduler provides a fully managed cron service that can trigger an HTTP or Pub/Sub target on a schedule, and Cloud Functions runs the Python code serverlessly with automatic retries and logging. This combination meets the requirement of no cluster or VM to maintain, supports retries on failure, and can publish to a Pub/Sub topic for notifications. It is the lightest managed option for a recurring single-step ETL job.

Why this answer

Cloud Scheduler plus Cloud Functions delivers a fully managed, serverless combination for a recurring ETL task. Cloud Scheduler handles the cron cadence and can target a function over HTTP or Pub/Sub, while Cloud Functions executes the Python code with automatic retries, integrated logging, and the ability to publish notifications. No cluster or VM is required, satisfying the operational constraints.

Exam trap

The trap here is treating Dataproc Serverless as a complete scheduling solution, when it only runs the compute and still needs an external trigger for a nightly cadence.

484
MCQmedium

A team notices that the latency for online predictions from a Vertex AI endpoint has increased significantly over the past hour. The model is a large TensorFlow model deployed with automatic scaling (minReplicaCount=2, maxReplicaCount=10). The CPU utilization of the deployed instances is consistently above 85%. What is the most likely cause of the increased latency?

A.The network latency between the client and the endpoint has increased due to regional issues.
B.The model is deployed with GPU acceleration, but the instances are using incorrect CUDA drivers.
C.The model is too large for the instance memory, causing disk swapping.
D.The model is CPU-bound, and the current replicas are saturated, causing queuing.
AnswerD

Sustained CPU above 85% across replicas indicates the model is compute-bound; requests queue behind in-flight inferences, inflating latency. Automatic scaling adds replicas but each new instance still saturates, so the bottleneck is CPU capacity, not network or model size.

Why this answer

The consistently high CPU utilization (above 85%) indicates that the existing replicas are saturated, unable to process incoming requests quickly enough. When all replicas are busy, new requests are queued, which directly increases latency. Automatic scaling can add more replicas up to maxReplicaCount=10, but if the scaling is slow or the traffic spike is sudden, queuing occurs first, causing the observed latency increase.

Exam trap

Google Cloud often tests the distinction between symptoms of CPU saturation (queuing/latency) versus memory or GPU issues; the trap here is that candidates may incorrectly attribute latency to network or hardware driver problems when the clear indicator is sustained high CPU utilization on existing instances.

How to eliminate wrong answers

Option A is wrong because network latency between client and endpoint is not indicated by CPU utilization of deployed instances; regional network issues would affect all requests uniformly, not correlate with high CPU. Option B is wrong because incorrect CUDA drivers would cause GPU-related errors or failures, not consistently high CPU utilization; the model would likely fail to run or produce errors, not just increase latency. Option C is wrong because disk swapping due to insufficient memory would manifest as high disk I/O and memory pressure, not primarily high CPU utilization; the symptom described is CPU-bound, not memory-bound.

485
MCQeasy

A company wants to implement a data lake on Google Cloud to store raw sensor data (unstructured binary files) and allow data scientists to run SQL queries on processed data. They expect to store terabytes of data and have different access patterns. Which combination of GCP services best meets these requirements?

A.Bigtable for raw data and Cloud Spanner for processed data
B.Cloud Storage for both raw and processed data
C.Cloud SQL for raw data and Cloud Dataproc for processing
D.Cloud Storage for raw data and BigQuery for processed data
AnswerD

Cloud Storage holds the raw unstructured binary sensor files cost-effectively at terabyte scale, while BigQuery provides serverless SQL analytics over the processed data. This pairing separates cheap object storage from query compute, matching the differing access patterns described.

Why this answer

Cloud Storage is the ideal service for storing raw, unstructured binary sensor data at petabyte scale, offering low-cost, durable object storage with multiple access tiers. BigQuery is a serverless, highly scalable data warehouse that allows data scientists to run SQL queries on processed data, with features like columnar storage and automatic optimization for analytical workloads. This combination directly addresses the need for raw storage and SQL-based analytics on processed data.

Exam trap

Google Cloud often tests the misconception that Cloud Storage can serve as a queryable database for SQL, when in fact it requires an external query engine like BigQuery or Dataproc for SQL access.

How to eliminate wrong answers

Option A is wrong because Bigtable is a NoSQL wide-column database optimized for real-time, low-latency access, not for storing raw unstructured binary files, and Cloud Spanner is a globally distributed relational database for transactional workloads, not for analytical SQL queries on processed data. Option B is wrong because while Cloud Storage can store both raw and processed data, it does not natively support SQL queries; data scientists would need an additional service like BigQuery or Dataproc to run SQL. Option C is wrong because Cloud SQL is a relational database for structured data, not designed for raw unstructured binary files, and Cloud Dataproc is a managed Spark/Hadoop service for processing, not a SQL query engine for processed data.

486
MCQeasy

A data engineer wants to automatically move objects from Standard storage class to Nearline after 30 days, and then to Archive after 365 days. Which Cloud Storage feature should they configure?

A.Object Versioning
B.Retention Policy
C.Bucket Lock
D.Object Lifecycle rule with SetStorageClass actions
AnswerD

An Object Lifecycle rule with SetStorageClass actions transitions objects automatically from Standard to Nearline at 30 days and to Archive at 365 days, satisfying both age-based tiering constraints. Lifecycle rules act on object age without manual intervention or application code.

Why this answer

Object Lifecycle rules in Google Cloud Storage allow you to automatically transition objects between storage classes (e.g., from Standard to Nearline after 30 days, then to Archive after 365 days) using the SetStorageClass action. This feature is specifically designed for automated lifecycle management, including deletion and class transitions, based on object age or other conditions.

Exam trap

Google often tests the distinction between lifecycle management (which changes storage classes) and retention/versioning features (which protect data but do not automate class transitions), leading candidates to confuse Object Versioning or Retention Policy with lifecycle rules.

How to eliminate wrong answers

Option A is wrong because Object Versioning is a feature that preserves non-current object versions to protect against accidental deletion or overwriting; it does not automate storage class transitions. Option B is wrong because Retention Policy is used to enforce a minimum retention period on objects, preventing deletion or modification, but it cannot change storage classes over time. Option C is wrong because Bucket Lock is a mechanism to permanently lock a retention policy, making it immutable; it does not provide any lifecycle-based storage class transitions.

487
MCQmedium

You manage a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must be updated to add a new transformation that enriches each message with data from Cloud SQL. You need to minimize downtime and ensure the pipeline continues processing without data loss. What should you do?

A.Use Cloud Scheduler to trigger a Cloud Function that updates the pipeline code in place.
B.Create a new Pub/Sub subscription, update the pipeline to read from it, and deploy a new pipeline.
C.Stop the existing pipeline, update the code, and start a new pipeline from the same subscription.
D.Use the Dataflow update feature to replace the pipeline with the new code while preserving the pipeline state.
AnswerD

The Dataflow update feature allows you to replace the running pipeline with a new version while maintaining the pipeline's state, such as Pub/Sub checkpoints and BigQuery write buffers. This minimizes downtime and ensures no data loss. It is the recommended approach for updating streaming pipelines in production when the update is compatible with the existing pipeline structure.

Why this answer

The Dataflow update feature is designed to replace a running pipeline with a new version while preserving state, enabling seamless updates with minimal downtime and no data loss. Stopping and restarting, creating new subscriptions, or using external triggers do not maintain pipeline state and risk data duplication or loss.

Exam trap

The trap here is assuming that stopping and restarting a pipeline from the same subscription is safe, but it can lead to duplicate or lost messages without proper checkpointing.

488
MCQeasy

You need to load a large CSV file from Cloud Storage into BigQuery. The file contains a header row and is comma-delimited. You want to ensure that the header row is skipped and that the schema is automatically detected. Which BigQuery load option should you use?

A.Set skip_leading_rows to 1 and autodetect to true.
B.Set allow_jagged_rows to true and ignore_unknown_values to true.
C.Set field_delimiter to '\t' and skip_leading_rows to 0.
D.Set max_bad_records to 1 and autodetect to false.
AnswerA

Setting skip_leading_rows to 1 tells BigQuery to ignore the first row (the header). Enabling autodetect allows BigQuery to infer the schema from the data. This combination is ideal for CSV files with headers when you don't want to manually define the schema. It is a common practice for loading external data efficiently.

Why this answer

To load a CSV with a header row and automatically detect the schema, you should use skip_leading_rows=1 to skip the header and autodetect=true to infer the schema. These options are part of the load job configuration in BigQuery. They streamline the loading process and reduce manual schema definition.

Exam trap

The trap here is confusing options that handle malformed data with those that manage headers and schema detection.

489
MCQmedium

A company runs a critical batch pipeline using Cloud Dataflow. The pipeline processes financial transactions and runs every hour. Recently, some runs have failed due to transient errors (e.g., network timeouts). The engineer wants to automatically retry failed runs without manual intervention. The pipeline is launched from a Cloud Composer DAG using DataflowPythonOperator. What is the BEST way to handle retries?

A.Add a DataflowJobStatusSensor in the DAG that waits for job completion and retries if failed.
B.Set the 'retries' parameter in the DAG's default_args to a positive integer.
C.Configure the Dataflow pipeline to automatically retry on failure using the --numberOfWorkerHarnessThreads option.
D.Use a Cloud Function triggered by Cloud Scheduler to re-launch the pipeline if the Dataflow job fails.
AnswerB

Setting retries in the DAG's default_args makes Cloud Composer automatically re-run the DataflowPythonOperator task after transient failures such as network timeouts, without manual intervention. Airflow's task-level retry mechanism satisfies the requirement for unattended recovery of hourly runs.

Why this answer

Cloud Composer (Apache Airflow) natively supports task-level retries via the 'retries' parameter in default_args. When a DataflowPythonOperator fails due to a transient error, Airflow automatically re-executes the task up to the specified number of retries, without requiring custom sensors or external triggers. This is the simplest and most reliable mechanism for handling transient failures in a DAG-driven pipeline.

Exam trap

The trap here is that candidates confuse Dataflow-level retry options (like --maxRetryAttempts) with Airflow task-level retries, or assume that a sensor or external trigger is required to detect and retry failures, when in fact Airflow's native retry parameter is the simplest and most appropriate solution for transient errors in a DAG-managed pipeline.

How to eliminate wrong answers

Option A is wrong because a DataflowJobStatusSensor only monitors job status and does not automatically retry the pipeline; it would require additional branching logic to relaunch the job, adding unnecessary complexity. Option C is wrong because --numberOfWorkerHarnessThreads controls parallelism within the Dataflow worker, not retry behavior on pipeline failure; retries are configured via --maxRetryAttempts or similar Dataflow pipeline options, not this flag. Option D is wrong because using a Cloud Function and Cloud Scheduler introduces an external dependency and latency, whereas Airflow's built-in retry mechanism is more direct and integrated with the DAG's execution context.

490
Multi-Selectmedium

Which TWO actions are recommended to improve the reliability of a Cloud Dataflow streaming pipeline that processes event data from Pub/Sub?

Select 2 answers
A.Use a pull subscription with a 10-second acknowledgment deadline.
B.Enable Dataflow Streaming Engine.
C.Enable exactly-once processing sinks (e.g., BigQuery with guaranteed row-level insertion).
D.Disable autoscaling to prevent worker churn.
E.Use micro-batch processing with a small batch size.
AnswersB, C

Streaming Engine moves pipeline state and shuffle execution off the worker VMs to the Dataflow service, so workers no longer persist state to attached disks. This reduces worker restart and pipeline update disruption, improving reliability for Pub/Sub streaming pipelines.

Why this answer

Option B is correct because Dataflow Streaming Engine moves pipeline state and shuffle execution out of the worker VMs and into the Dataflow service, which improves reliability by decoupling state management from worker lifecycle events such as restarts, autoscaling, and updates. Option C is correct because enabling exactly-once processing sinks (for example, BigQuery with row-level deduplication via insertion IDs) prevents duplicate or lost records when retries occur, which is essential for reliable event processing from Pub/Sub. Option A is not recommended because a 10-second acknowledgment deadline is too short for many streaming workloads and would cause premature redelivery and duplicate processing rather than improving reliability.

Option D is incorrect because disabling autoscaling does not improve reliability and can actually reduce throughput and resilience during load spikes. Option E is incorrect because micro-batching with a small batch size is not a Dataflow reliability best practice and can increase overhead and latency without addressing state or delivery guarantees.

Exam trap

The trap here is that candidates often confuse reliability with throughput or latency, and may incorrectly choose micro-batching or disabling autoscaling as reliability improvements, when in fact Dataflow's reliability comes from its managed backend services like Streaming Engine.

491
MCQeasy

An organization uses BigQuery on-demand pricing. To control costs, they want to estimate the bytes processed by a query before running it. Which command or method should they use?

A.Use the bq query --dry_run command
B.Use bq ls to list table sizes
C.Use BigQuery reservations to get cost estimate
D.Use INFORMATION_SCHEMA.JOBS_BY_PROJECT to view past costs
AnswerA

The `bq query --dry_run` flag validates a query and returns the bytes it would process without executing it, so no on-demand charges are incurred. This directly satisfies the requirement to estimate bytes processed beforehand, letting the organisation predict cost before committing to the query.

Why this answer

The bq query --dry_run flag validates a query and returns the estimated bytes that would be processed without actually executing it or incurring charges. This is the standard method to estimate on-demand BigQuery costs before running a query, since on-demand pricing is based on bytes processed.

Exam trap

PDE often tests the difference between pre-execution estimation and post-hoc cost inspection, so the trap is choosing INFORMATION_SCHEMA.JOBS_BY_PROJECT (which shows past bytes billed) instead of --dry_run for estimating before running.

How to eliminate wrong answers

Option B is wrong because bq ls only lists datasets, tables, or jobs—it does not estimate query bytes processed. Option C is wrong because reservations are for capacity-based (flat-rate) pricing and do not provide a per-query byte estimate. Option D is wrong because INFORMATION_SCHEMA.JOBS_BY_PROJECT shows historical job metadata and bytes billed after the fact, not a pre-execution estimate.

492
MCQmedium

A data engineer is designing a BigQuery table to store e-commerce order events. Each row contains an order_id (STRING, high cardinality), customer_id (STRING, high cardinality), order_timestamp (TIMESTAMP), and order_amount (NUMERIC). Queries frequently filter on order_id for lookups and also scan by order_timestamp ranges for daily reporting. The engineer wants to minimize bytes scanned by both query patterns. What should the engineer do?

A.Partition the table by order_timestamp and cluster by order_id.
B.Partition the table by order_id and cluster by order_timestamp.
C.Create a materialized view that pre-aggregates daily order totals and query the view instead.
D.Enable BigQuery BI Engine on the table and rely on in-memory caching for all queries.
AnswerA

Partitioning by order_timestamp limits daily range scans to the relevant date partitions, and clustering by order_id sorts data within each partition so point lookups on order_id read only the matching blocks. This combination addresses both access patterns without duplicating data or adding maintenance overhead.

Why this answer

Partitioning by order_timestamp restricts daily range scans to the relevant date partitions, while clustering by order_id sorts rows within each partition so that point lookups read only matching blocks. This design targets both query patterns and minimizes bytes scanned without adding a separate pipeline or storage layer.

Exam trap

The trap here is assuming clustering alone can replace partitioning for timestamp range filters, when clustering only helps when the filter columns are a prefix of the cluster keys and does not prune whole partitions.

493
MCQmedium

A financial analytics team ingests trade records into BigQuery every minute. Queries filter almost exclusively on trade_date and account_id, and the table grows by roughly 400 GB per day. Analysts report that monthly reports scanning the last 30 days are slow and expensive. The team wants to reduce bytes scanned without changing the ingestion pipeline. What should they do?

A.Enable BigQuery BI Engine reservation for the reporting dataset.
B.Partition the table by trade_date and cluster it by account_id.
C.Convert the table to a BigQuery external table over Cloud Storage Parquet files.
D.Create a materialized view that pre-aggregates daily totals per account.
AnswerB

Partitioning by trade_date lets BigQuery prune all partitions outside the 30-day window, and clustering by account_id keeps account-filtered blocks co-located within each partition. Together they cut bytes scanned for both the date range and the account predicate, directly addressing the slow, expensive monthly reports without touching ingestion.

Why this answer

Partitioning by the trade_date column allows BigQuery to skip partitions outside the query's date range, and clustering by account_id further prunes blocks within each retained partition. This combination directly reduces bytes scanned for the described monthly reports, improving both latency and cost while leaving the ingestion pipeline untouched.

Exam trap

The trap here is assuming a materialized view or BI Engine is a substitute for partitioning and clustering when the real issue is scanning too much table data.

494
MCQmedium

You need to create a BigQuery table that stores customer transaction data. The table will be queried frequently by a customer_id column to retrieve recent transactions (last 30 days). Which table design optimizes query performance and cost?

A.Partition by customer_id and cluster by transaction_date
B.Partition by ingestion_time and cluster by customer_id
C.Cluster by transaction_date and customer_id without partitioning
D.Partition by transaction_date and cluster by customer_id
AnswerD

Partitioning on transaction_date lets queries prune to the last 30 days, and clustering on customer_id sorts data within each partition so lookups scan far fewer bytes. Together they cut both query latency and bytes-billed cost for the stated access pattern.

Why this answer

Partitioning by transaction_date aligns with the query filter (last 30 days), so BigQuery prunes all partitions outside that window, scanning only the relevant date range. Clustering by customer_id then sorts data within each partition, so the customer_id filter benefits from block-level pruning. This combination minimizes both bytes scanned (cost) and query latency.

Exam trap

PDE often tests the misconception that the most selective column (customer_id) should be the partition key, when the correct rule is that the partition column must match the time-based filter predicate to enable pruning.

How to eliminate wrong answers

Option A is wrong because partitioning by customer_id creates one partition per customer, which explodes partition count (BigQuery's practical limit is ~4,000 partitions per table) and does not help filter by date range. Option B is wrong because partitioning by ingestion_time does not match the transaction_date filter, so no partition pruning occurs for 'last 30 days' queries. Option C is wrong because clustering alone provides no partition pruning — the entire table is scanned and only block-level pruning applies, which is far less efficient than partitioning plus clustering.

495
MCQhard

A company uses Vertex AI Feature Store for serving features. They have a high-throughput online serving requirement. Which configuration should they use?

A.Cloud Storage with high-memory instances
B.Bigtable as serving source
C.Firestore
D.Vertex AI Feature Store with online serving enabled
AnswerD

Online serving must be explicitly enabled on the Vertex AI Feature Store instance to expose low-latency feature retrieval through the online store. Without it, only batch serving is available, which cannot meet the high-throughput online serving requirement.

Why this answer

Vertex AI Feature Store with online serving enabled is the correct choice because it is specifically designed for low-latency, high-throughput retrieval of feature values for online predictions. It uses a managed Bigtable backend optimized for real-time serving, ensuring consistent performance under high request loads without requiring manual infrastructure management.

Exam trap

Google Cloud often tests the misconception that any low-latency database (like Bigtable or Firestore) can directly replace Vertex AI Feature Store, ignoring the managed orchestration, feature registry, and point-in-time lookup capabilities that are essential for consistent online serving in ML workflows.

How to eliminate wrong answers

Option A is wrong because Cloud Storage is a blob storage service with high latency and no indexing for real-time feature lookups, making it unsuitable for high-throughput online serving. Option B is wrong because Bigtable is a NoSQL database that can serve features, but it requires manual configuration, scaling, and integration with Vertex AI Feature Store, whereas the Feature Store provides a managed, optimized serving layer with built-in consistency and monitoring. Option C is wrong because Firestore is a document database designed for mobile and web apps with moderate throughput, not for the sub-millisecond latency and high concurrency required by ML feature serving at scale.

496
MCQmedium

You are designing a streaming pipeline that needs to handle sudden spikes in traffic without losing data. The pipeline uses Pub/Sub and Dataflow. Which configuration ensures data is not lost if Dataflow falls behind?

A.Use Pub/Sub with a pull subscription and set the message retention duration to 7 days
B.Use Cloud Pub/Sub Lite with a smaller retention period
C.Use Pub/Sub with a push subscription and increase the acknowledgment deadline
D.Use Pub/Sub with exactly-once delivery and Dataflow with at-least-once processing
AnswerA

A pull subscription with seven-day retention keeps unacknowledged messages durably stored while Dataflow's backlog grows during spikes, so no data is dropped. This satisfies the stem's no-loss constraint when the pipeline falls behind, since push subscriptions cannot replay expired messages.

Why this answer

A Pub/Sub pull subscription with a 7-day message retention duration ensures that if Dataflow falls behind, messages are retained in the subscription backlog for up to 7 days, preventing data loss. Dataflow pulls messages and acknowledges them only after processing, so unacknowledged messages remain available for redelivery. This configuration decouples ingestion from processing and provides a buffer for spikes.

Exam trap

The trap is assuming that exactly-once delivery or push subscriptions prevent data loss, when actually retention and pull-based backpressure are what protect against Dataflow lag.

How to eliminate wrong answers

Option B is wrong because Pub/Sub Lite has a smaller retention period (up to 7 days but typically less) and is not designed for the same durability guarantees; it also requires manual capacity management. Option C is wrong because push subscriptions push messages to an endpoint, and increasing the acknowledgment deadline only extends the time before redelivery, but if Dataflow is overwhelmed, messages may still be lost or the endpoint may fail. Option D is wrong because exactly-once delivery in Pub/Sub does not guarantee no data loss if Dataflow falls behind; it only prevents duplicates, and at-least-once processing in Dataflow may still drop messages if the pipeline is overwhelmed.

497
MCQeasy

A financial services firm needs to design a batch processing system on Google Cloud to analyze large volumes of historical transaction data stored in Cloud Storage. The data is in Parquet format and must be processed using Apache Spark. The firm wants to minimize operational overhead and only pay for the resources used during job execution. Which Google Cloud service should they use?

A.Cloud Dataflow with a custom Apache Spark runner
B.Cloud Dataproc Serverless for Spark
C.BigQuery with external tables over Cloud Storage
D.Cloud Dataproc with a long-running cluster
AnswerB

Cloud Dataproc Serverless for Spark is a fully managed, serverless Spark environment that eliminates the need to provision or manage clusters. It automatically scales resources based on workload and charges only for the resources consumed during job execution. It natively supports Spark jobs and can read Parquet data from Cloud Storage, making it the ideal choice for minimizing operational overhead and paying only for what is used.

Why this answer

Cloud Dataproc Serverless for Spark provides a fully managed Spark environment without cluster provisioning. It charges only for the resources used during job execution, aligning with the cost and operational requirements. It supports reading Parquet from Cloud Storage and is designed for batch processing workloads, making it the best fit for analyzing historical transaction data with Spark while minimizing overhead.

Exam trap

The trap here is confusing Dataflow with a Spark execution service or assuming that Dataproc requires a persistent cluster, when serverless Spark is available.

498
MCQmedium

You have a BigQuery table that is used by multiple teams. To save costs, you want to provide a consistent view of the data as of a specific point in time without creating full copies. Which BigQuery feature should you use?

A.Authorized views
B.Materialized views
C.Table snapshot
D.Table clone
AnswerC

A table snapshot captures the table's data at a specific point in time using a copy-on-write reference, so it is queryable and cheap to create without duplicating storage. This gives multiple teams a consistent historical view while avoiding the cost of full table copies.

Why this answer

A BigQuery table snapshot creates a lightweight, read-only copy of a table at a specific point in time without duplicating the underlying data, and it is the correct feature for providing a consistent point-in-time view cheaply. Snapshots preserve the table state at creation and can be queried or restored, making them ideal for this requirement.

Exam trap

PDE often tests the difference between snapshots (read-only, point-in-time, copy-on-write) and clones (writable, diverging), so the trap is choosing table clone or materialized view when the requirement is a frozen, consistent point-in-time view.

How to eliminate wrong answers

Option A is wrong because authorized views control access to data, not point-in-time consistency, and do not create a snapshot of the data state. Option B is wrong because materialized views precompute and cache query results for acceleration, not for preserving a point-in-time copy, and they refresh as the base table changes. Option D is wrong because a table clone creates a writable copy that shares storage initially but diverges on writes and does not freeze a point-in-time state.

499
MCQhard

A company needs to continuously synchronize customer data changes from an on-premises Oracle database to BigQuery for near-real-time analytics. The Oracle database has Change Data Capture (CDC) enabled. Which Google Cloud service should be used to stream these changes with minimal latency and schema evolution support?

A.Deploy a Dataflow pipeline with a JDBC source and Pub/Sub
B.Use Cloud SQL with a read replica and enable binary logging
C.Use Transfer Appliance to copy Oracle data periodically
D.Use Datastream to stream CDC changes from Oracle to BigQuery
AnswerD

Datastream provides native Oracle CDC replication, reading redo logs to stream changes into BigQuery with minimal latency. Its built-in schema evolution handling automatically propagates source DDL changes, satisfying the stem's requirement for continuous near-real-time synchronisation without custom extraction code.

Why this answer

Datastream is Google Cloud's fully managed, serverless CDC and replication service that natively supports Oracle sources (via LogMiner or XStream) and BigQuery destinations. It streams changes with low latency, handles schema evolution automatically, and requires no pipeline code to maintain. This matches the requirement for minimal latency and schema evolution support.

Exam trap

PDE often tests whether candidates recognize that Datastream is the purpose-built managed CDC service, while Dataflow+JDBC is a custom polling approach that candidates mistakenly equate with CDC.

How to eliminate wrong answers

Option A is wrong because a custom Dataflow JDBC pipeline polls the source rather than reading CDC logs, introducing latency and requiring custom code to handle schema changes — it does not leverage Oracle's CDC capability. Option B is wrong because Cloud SQL is a managed MySQL/PostgreSQL/SQL Server service; it cannot act as a read replica of an on-premises Oracle database, and binary logging is a MySQL concept, not Oracle. Option C is wrong because Transfer Appliance is a physical appliance for bulk offline data transfer, which is the opposite of near-real-time streaming and provides no CDC or schema evolution.

500
Multi-Selecthard

A data engineer is designing a Cloud Storage layout for a data lake that will be queried by BigQuery external tables and by Dataproc jobs. The engineer wants to minimize query cost and improve scan performance across both engines. Which two practices should the engineer follow? (Choose two.)

Select 2 answers
A.Enable object versioning on the bucket to preserve historical query results.
B.Partition the data in Cloud Storage using a Hive-style key prefix such as dt=YYYY-MM-DD.
C.Store each row as a separate object to maximize parallelism during scans.
D.Store data in columnar formats such as Parquet or ORC instead of CSV.
E.Use the Standard storage class for all objects to ensure the lowest possible access latency.
AnswersB, D

Hive-style partitioning lets BigQuery external tables and Dataproc prune entire prefixes when a query filters on the partition column, so only relevant directories are listed and read. This reduces both listing overhead and bytes scanned, improving performance and lowering cost for both engines.

Why this answer

Columnar formats let both engines read only referenced columns and compress data efficiently, and Hive-style partitioning lets them prune entire prefixes when filters match the partition key. Together these reduce bytes scanned and listing overhead, which lowers query cost and improves scan performance for BigQuery external tables and Dataproc alike.

Exam trap

The trap here is focusing on storage-class or versioning settings, which affect durability and retrieval cost, instead of the data layout choices that actually reduce bytes scanned.

501
Multi-Selecteasy

You need to deploy a reusable Dataflow pipeline that can be executed with different parameters from Cloud Composer. Which TWO components should you use? (Choose 2)

Select 2 answers
A.Direct runner
B.Dataflow Flex Template
C.Cloud Composer with DataflowStartFlexTemplateOperator
D.Dataflow Classic Template
E.Cloud Scheduler
AnswersB, C

A Flex Template packages the pipeline as a Docker image with a metadata parameter spec, so the same artefact runs repeatedly with different runtime parameters. This satisfies the reusability constraint, unlike classic templates, which require rebuilding or redeploying the pipeline for parameter changes.

Why this answer

Option B (Dataflow Flex Template) is correct because Flex Templates package the pipeline as a Docker image plus a template spec file, allowing the same pipeline to be reused and launched with different runtime parameters (such as input/output locations) via the templates launch API. Option C (Cloud Composer with DataflowStartFlexTemplateOperator) is correct because this operator is purpose-built to submit a Flex Template job from an Airflow/Composer DAG, passing the required parameters and letting Composer orchestrate the pipeline execution. Option A (Direct runner) is incorrect because it runs the pipeline locally for testing rather than deploying a reusable job on Dataflow.

Option D (Dataflow Classic Template) is not the best fit here since Classic Templates are less flexible for parameterization and dependency packaging compared with Flex Templates. Option E (Cloud Scheduler) is incorrect because it only triggers jobs on a schedule and does not deploy or parameterize a Dataflow pipeline.

Exam trap

PDE often tests the distinction between Flex Templates and Classic Templates, as candidates may choose Classic Templates for reusability but overlook the need for runtime parameterization and the specific operator for Cloud Composer integration.

502
MCQeasy

A data engineer needs to ingest on-premises Oracle CDC data into BigQuery in near real-time with minimal operational overhead. Which service should they use?

A.Pub/Sub + Dataflow
B.Storage Transfer Service
C.Transfer Appliance
D.Datastream
AnswerD

Datastream provides serverless change data capture from Oracle into BigQuery with minimal operational overhead, replicating changes in near real-time. It satisfies both constraints directly, unlike batch extraction tools or custom pipelines that require managing infrastructure.

Why this answer

Datastream is purpose-built for streaming change data capture (CDC) from Oracle and other sources into BigQuery with near-real-time latency and minimal operational overhead. It handles schema propagation, checkpointing, and automatic retries, eliminating the need to manage custom ingestion pipelines.

Exam trap

Google often tests the distinction between batch migration tools (Storage Transfer Service, Transfer Appliance) and streaming CDC services (Datastream), leading candidates to choose a batch option when the question explicitly requires near-real-time ingestion.

How to eliminate wrong answers

Option A is wrong because Pub/Sub + Dataflow requires building and maintaining a custom pipeline to handle Oracle CDC, including log mining and transformation logic, which increases operational overhead compared to a managed service. Option B is wrong because Storage Transfer Service is designed for bulk batch transfers of files from cloud or on-premises storage to Google Cloud, not for streaming CDC from a live database. Option C is wrong because Transfer Appliance is a physical device for offline, high-volume data migration, which cannot provide near-real-time streaming and introduces significant latency.

503
MCQmedium

A financial analytics team uses Looker to explore BigQuery data. They need to allow business users to filter by a custom date range that is not tied to an existing dimension. The date range must be user-input at query time. What is the best approach in Looker?

A.Create an explore with a custom filter field in the Looker UI
B.Use a filter parameter directly on the date dimension
C.Add a dimension with a yesno filter that toggles the date range
D.Create a parameter in LookML using Liquid templating
AnswerD

A LookML parameter with Liquid templating injects a user-supplied value into the generated SQL at query time, letting business users enter an arbitrary date range not bound to any existing dimension. This satisfies the stem's user-input-at-query-time constraint.

Why this answer

Creating a parameter in LookML using Liquid templating allows business users to input a custom date range at query time. Parameters are user-input fields that can be referenced in SQL queries via Liquid, enabling dynamic filtering. This is the most flexible approach for ad-hoc date ranges not tied to existing dimensions.

Exam trap

PDE often tests Looker customization, and candidates might think UI-based filters suffice, but for arbitrary user input, LookML parameters with Liquid are required.

How to eliminate wrong answers

Option A is wrong because a custom filter field in the Looker UI is typically based on existing dimensions and cannot easily accept arbitrary date ranges without a parameter. Option B is wrong because a filter parameter directly on a date dimension is not a standard Looker feature; parameters are defined in LookML. Option C is wrong because a yesno filter toggles a predefined condition, not a custom date range.

504
MCQmedium

An organization needs to trigger a Cloud Run service whenever a new file is uploaded to a specific Cloud Storage bucket. Which service should they use to set up this event-driven architecture?

A.Eventarc with a trigger for Cloud Storage events
B.Pub/Sub notifications on the bucket with a push subscription to Cloud Run
C.Cloud Scheduler calling Cloud Run on a schedule
D.Cloud Functions with a GCS trigger
AnswerA

Eventarc natively routes Cloud Storage object-finalise events to Cloud Run, satisfying the requirement to trigger on each new upload. It delivers events through the CloudEvents standard with built-in filtering by bucket and event type, so no polling or manual Pub/Sub plumbing is needed.

Why this answer

Eventarc is the recommended Google Cloud service for routing events from Cloud Storage (and 90+ other sources) to Cloud Run, Cloud Functions, or GKE. It provides a native Cloud Storage trigger type, handles authentication, retries, and dead-lettering, and integrates directly with Cloud Run's event delivery model. This is the canonical event-driven pattern for GCS-to-Cloud-Run.

Exam trap

PDE often tests the distinction between Eventarc (managed event routing) and raw Pub/Sub notifications — candidates pick Pub/Sub because it 'works,' missing that Eventarc is the recommended, lower-overhead abstraction for Cloud Run triggers.

How to eliminate wrong answers

Option B is wrong because while Pub/Sub notifications on a bucket with a push subscription to Cloud Run can technically work, it requires manual Pub/Sub topic creation, IAM wiring, and does not provide the first-class Cloud Storage event schema or filtering that Eventarc offers — it is more operational overhead and not the recommended pattern. Option C is wrong because Cloud Scheduler is time-based, not event-based; it cannot react to a file upload. Option D is wrong because Cloud Functions with a GCS trigger targets Cloud Functions (1st/2nd gen), not Cloud Run — the question specifically asks about triggering a Cloud Run service.

505
MCQhard

You manage a Cloud Composer environment that runs a critical DAG every hour. The DAG includes a task that calls a Cloud Function to process data. Recently, the Cloud Function started taking longer than expected, causing the DAG to exceed its SLA. You need to detect this delay and automatically retry the task if it fails due to timeout, while minimizing changes to the DAG. What should you do?

A.Use a Cloud Monitoring alert on the Cloud Function's execution time and manually trigger the DAG upon alert.
B.Increase the Cloud Function's timeout setting to the maximum allowed and rely on the DAG's default retry behavior.
C.Set the task's execution_timeout parameter to a value slightly above the expected runtime and configure retries with a delay.
D.Add a Python operator that polls the Cloud Function's status and raises an exception if it exceeds a threshold, then set retries on that operator.
AnswerC

The execution_timeout parameter defines the maximum time a task can run before it is killed and marked as failed. Setting it appropriately allows the task to be retried if it exceeds the timeout. Configuring retries ensures automatic retry on failure. This approach requires minimal changes to the DAG and directly addresses the timeout issue, making it the most efficient solution.

Why this answer

Using the task's execution_timeout parameter allows Airflow to enforce a time limit on the task. If the task exceeds this limit, it fails and can be automatically retried based on the retries and retry_delay settings. This requires only a small change to the DAG's task definition and leverages native Airflow features, effectively addressing both detection of delays and automatic retries.

Exam trap

The trap here is focusing on the Cloud Function's timeout instead of the Airflow task's execution_timeout, which controls retries at the orchestration level.

506
Multi-Selectmedium

A company is building a real-time anomaly detection pipeline using Dataflow. Events are ingested from Pub/Sub, and the pipeline must compute a sliding window average every minute over a 1-hour window. Which TWO configurations are required for this pipeline? (Choose 2)

Select 2 answers
A.Set the pipeline to use event time for watermarking.
B.Use a Sliding window of 1 hour with a 1-minute slide.
C.Use a Fixed window of 1 minute.
D.Use stateful processing with a custom timer.
E.Set the pipeline to use processing time for watermarking.
AnswersA, B

Event-time watermarking derives progress from the timestamp embedded in each Pub/Sub message rather than arrival time, so the one-hour sliding window aggregates correctly despite network delays and out-of-order events. Without it, window results would be non-deterministic and inaccurate.

Why this answer

Option A is correct because a sliding window over event data must be based on event time so that late-arriving events are assigned to the correct window and watermarks track the progress of event time rather than the wall-clock time at which elements are processed. Option B is correct because the requirement is a 1-hour window that is recomputed every minute, which is exactly a SlidingWindow with a duration of 1 hour and a period (slide) of 1 minute. Option C is incorrect because a FixedWindow of 1 minute produces independent non-overlapping 1-minute aggregates, not a 1-hour average refreshed each minute.

Option D is incorrect because stateful processing with a custom timer is a lower-level mechanism and is not required when Apache Beam's built-in sliding windows already express the desired semantics. Option E is incorrect because processing-time watermarking ignores the event timestamps and would misassign late or out-of-order Pub/Sub messages, breaking the anomaly detection logic.

Exam trap

The trap is assuming a FixedWindow can produce a rolling average — candidates who conflate 'window size' with 'update frequency' pick FixedWindow of 1 minute and miss that sliding windows require both a duration and a period.

507
MCQhard

Refer to the exhibit. A BigQuery dataset is shared with the group 'analysts@example.com' using the IAM policy shown. A user who is a member of this group reports that they cannot run queries on the dataset, though they can see the tables. What is the most likely reason?

A.The group needs the 'roles/bigquery.jobUser' role at the project level.
B.The user is using an incorrect client library version.
C.The user's account is not activated in the group membership.
D.The dataset has an organization policy that denies query access.
AnswerA

BigQuery separates metadata visibility from query execution: dataset-level roles such as `roles/bigquery.dataViewer` let the user list tables, but running jobs requires `roles/bigquery.jobUser` granted at the project level. Because the group only holds dataset-scoped access, the user can see tables yet cannot execute queries, satisfying the stem's constraint.

Why this answer

The IAM policy grants the 'roles/bigquery.dataViewer' role at the dataset level, which allows the user to see tables but not run queries. To run queries, the user also needs the 'roles/bigquery.jobUser' role at the project level, because BigQuery query jobs are project-scoped resources. Without this role, the user lacks permission to create query jobs, even though they can view dataset metadata.

Exam trap

This question tests the distinction between dataset-level and project-level roles in BigQuery. Candidates often incorrectly assume that dataset-level view permissions are sufficient to run queries, but query jobs require the 'roles/bigquery.jobUser' role at the project level.

How to eliminate wrong answers

Option B is wrong because client library version does not affect IAM permissions; authentication and authorization are handled by Google Cloud IAM, not by the library version. Option C is wrong because if the user's account were not activated in the group membership, they would not be able to see the tables at all, as the dataset-level view permission would not apply. Option D is wrong because an organization policy that denies query access would typically block all query operations for all users, not just this user, and the user can see tables, which contradicts a blanket deny on queries.

508
Multi-Selecteasy

A data engineering team is operationalizing a machine learning model for real-time fraud detection. The model must process transactions with sub-100ms latency and be highly available. Which TWO strategies should the team implement?

Select 2 answers
A.Deploy the model to multiple Google Cloud regions for failover.
B.Deploy the model to a single zone to minimize cross-zone latency.
C.Use Cloud Batch for asynchronous prediction.
D.Optimize the model by pruning or quantizing to reduce size.
E.Store the model in Cloud Storage and load it on each request.
AnswersA, D

Deploying across multiple Google Cloud regions directly satisfies the high-availability constraint by eliminating single-region failure as a single point of outage. Regional redundancy lets traffic fail over when one region degrades, preserving continuous fraud scoring. It does not by itself guarantee sub-100ms latency, which the second strategy must address.

Why this answer

Option A is correct because deploying the model to multiple Google Cloud regions provides geographic redundancy and failover, which directly supports the high-availability requirement for a real-time fraud detection service. Option D is correct because pruning or quantizing the model reduces its size and computational cost, which helps achieve the sub-100ms latency target for real-time inference. Option B is incorrect because a single zone creates a single point of failure and does not meet the high-availability requirement, even if it reduces cross-zone latency.

Option C is incorrect because Cloud Batch is designed for asynchronous, batch-oriented workloads, not real-time sub-100ms predictions. Option E is incorrect because loading the model from Cloud Storage on every request adds significant network and I/O latency, making the sub-100ms target impractical.

Exam trap

Google Cloud often tests the misconception that single-zone deployment minimizes latency, but the real trade-off is between availability and negligible intra-region latency, making multi-region deployment the correct choice for high availability.

509
MCQmedium

A company uses Google Ads and wants to automatically load their advertising data into BigQuery daily. They also need to transform the data with SQL and schedule a recurring query. Which combination of services meets these requirements with minimal operational overhead?

A.Cloud Functions triggered by Cloud Scheduler to call Google Ads API and load into BigQuery
B.Cloud Composer to extract Google Ads API and Dataflow to transform
C.Storage Transfer Service to move CSV files to GCS, then load into BigQuery
D.BigQuery Data Transfer Service for Google Ads and scheduled queries
AnswerD

BigQuery Data Transfer Service natively ingests Google Ads data on a daily schedule, and BigQuery scheduled queries run the SQL transformations recurringly. Together they satisfy the daily load, SQL transformation, and scheduling requirements with no servers or orchestration code to maintain.

Why this answer

BigQuery Data Transfer Service has a built-in Google Ads connector that automatically loads advertising data into BigQuery on a schedule, and BigQuery scheduled queries let you transform that data with SQL on a recurring basis — both fully managed with no infrastructure. This combination meets the daily load and SQL transformation requirements with minimal operational overhead.

Exam trap

PDE often tests whether candidates choose the fully managed native connector (BigQuery DTS) over custom code (Cloud Functions/Composer) when the question emphasizes 'minimal operational overhead.'

How to eliminate wrong answers

Option A is wrong because a Cloud Functions + Cloud Scheduler approach requires custom code to call the Google Ads API, handle pagination, retries, and schema mapping — significantly more operational overhead than the managed connector. Option B is wrong because Cloud Composer (managed Airflow) plus Dataflow is heavyweight for a simple daily load-and-transform; it introduces DAG maintenance, cluster costs, and pipeline code that the question explicitly wants to avoid. Option C is wrong because Storage Transfer Service moves files between storage systems; it does not extract data from the Google Ads API, and CSV-based loading adds manual steps and latency.

510
MCQmedium

A team uses Vertex AI AutoML Tables to train a model. They need to deploy the model for real-time predictions with high availability. Which deployment configuration should they use?

A.Export as a Cloud Function
B.Deploy to a Vertex AI Endpoint with 1 replica
C.Use a Vertex AI Batch Prediction job
D.Deploy to a Vertex AI Endpoint with multiple replicas and auto-scaling
AnswerD

Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling directly satisfies the high-availability requirement: replicas distribute traffic across instances, so a single node failure does not take the service down, while auto-scaling absorbs real-time prediction load spikes without manual intervention.

Why this answer

For real-time predictions with high availability, you need a deployment that can handle traffic spikes and failover. Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling ensures that the model is served from multiple instances, providing redundancy and the ability to scale up or down based on demand. This configuration meets the high-availability requirement by distributing load and automatically recovering from instance failures.

Exam trap

The trap here is that candidates often confuse batch prediction with real-time serving, or assume that a single replica is sufficient for high availability, not realizing that high availability requires redundancy and automatic scaling.

How to eliminate wrong answers

Option A is wrong because exporting as a Cloud Function is not a deployment method for Vertex AI AutoML Tables models; Cloud Functions are for serverless event-driven code, not for hosting ML model endpoints with real-time prediction capabilities. Option B is wrong because deploying to a Vertex AI Endpoint with only 1 replica provides no redundancy or high availability; if that single instance fails or becomes overloaded, predictions will be unavailable. Option C is wrong because a Vertex AI Batch Prediction job is designed for asynchronous, offline predictions on large datasets, not for real-time, low-latency serving.

511
MCQmedium

A financial services company receives real-time stock trade data via Pub/Sub. They need to enrich each trade with reference data from a Cloud SQL table and write the results to BigQuery for real-time analytics. The enrichment must handle late-arriving data and ensure exactly-once processing. Which Dataflow streaming pipeline configuration should be used?

A.Use a Dataflow Flex Template that reads from Pub/Sub, joins in memory, and writes to BigQuery using legacy streaming inserts
B.Use Pub/Sub to BigQuery template with streaming inserts and a side input from Cloud SQL
C.Build a custom Dataflow pipeline using the Storage Write API with exactly-once semantics and a side input from Cloud SQL
D.Deploy a Dataproc Spark Streaming job that reads from Pub/Sub, enriches via JDBC, and writes to BigQuery
AnswerC

The Storage Write API's exactly-once mode prevents duplicate writes to BigQuery, while a side input broadcasts the Cloud SQL reference table to workers for enrichment. Event-time windows with allowed lateness handle late-arriving trades, satisfying both stated requirements.

Why this answer

The BigQuery Storage Write API with exactly-once semantics is the only option that satisfies both the exactly-once requirement and the need for enrichment with Cloud SQL reference data. A custom Dataflow pipeline can use the Storage Write API's exactly-once mode while applying a side input (or CoGroupByKey) to join the streaming trades with the Cloud SQL reference table. This combination handles late-arriving data via allowed lateness and windowing while guaranteeing no duplicate writes to BigQuery.

Exam trap

PDE often tests the distinction between legacy streaming inserts (at-least-once, duplicates possible) and the Storage Write API's exactly-once mode — candidates who overlook the exactly-once requirement pick a template or legacy insert option.

How to eliminate wrong answers

Option A is wrong because legacy streaming inserts into BigQuery do not provide exactly-once semantics and can produce duplicate rows on retries. Option B is wrong because the Pub/Sub to BigQuery template uses streaming inserts (at-least-once) and does not natively support a Cloud SQL side input for enrichment. Option D is wrong because Dataproc Spark Streaming with JDBC enrichment adds operational complexity and does not inherently provide exactly-once writes to BigQuery, making it a poor fit for the stated requirements.

512
MCQeasy

A data engineer needs to choose a storage service for a new application that requires a schemaless document store with automatic multi-region replication, strong consistency for reads, and native mobile SDK support. The application must scale to millions of users without manual sharding. Which Google Cloud service should the engineer select?

A.Cloud Spanner
B.Cloud Bigtable
C.Cloud Firestore in Native mode
D.Cloud SQL for MySQL
AnswerC

Firestore in Native mode is a schemaless document database with automatic multi-region replication, strong consistency for document reads, and first-class mobile SDKs. It scales horizontally without manual sharding and is designed for exactly this class of mobile and web application.

Why this answer

Firestore in Native mode provides a schemaless document model, automatic multi-region replication, strong consistency for reads, and native mobile SDKs, and it scales horizontally without manual sharding. These characteristics match the application's requirements more closely than the relational or wide-column alternatives.

Exam trap

The trap here is conflating Spanner's global strong consistency with the document-store and mobile-SDK requirements, when Spanner is a relational engine with a fixed schema.

513
MCQmedium

Refer to the exhibit. A data scientist deploys a model using this configuration. Users report that after a few hours of inactivity, the first prediction request takes over 30 seconds. What is the most likely cause?

A.The automatic scaling configuration allows scaling down to zero replicas, causing a cold start on the first request.
B.The network latency between the client and the endpoint is high due to regional distance.
C.The endpoint is misconfigured with the wrong regional endpoint.
D.The model is too large and exceeds the instance memory.
AnswerA

Scaling to zero removes all serving containers, so the next request must provision a replica and load the TensorFlow model, producing a cold-start delay of tens of seconds. Keeping a minimum replica count avoids this.

Why this answer

The automatic scaling configuration that allows scaling down to zero replicas means that after a period of inactivity, all model replicas are terminated. When a new prediction request arrives, the endpoint must provision a new replica from scratch, which involves loading the model artifacts, initializing the inference container, and performing health checks. This cold start process typically takes 30 seconds or more, matching the reported behavior.

Exam trap

Google Cloud often tests the distinction between cold start latency (caused by scaling to zero) and persistent performance issues like network latency or resource exhaustion, so candidates must recognize that a delay only after inactivity points to replica provisioning, not a constant problem.

How to eliminate wrong answers

Option B is wrong because network latency due to regional distance would cause consistent high latency on every request, not just the first request after a period of inactivity. Option C is wrong because a misconfigured regional endpoint would result in persistent errors or high latency on all requests, not a delay only after inactivity. Option D is wrong because if the model exceeded instance memory, the endpoint would fail to serve predictions consistently or return out-of-memory errors, not exhibit a delay only on the first request after inactivity.

514
MCQhard

A data scientist is training a binary classification model on an imbalanced dataset (95% negative, 5% positive) using AutoML Tables. Which strategy should they use to handle the class imbalance?

A.Set the budget to a higher value to allow more training on minority class.
B.Use SMOTE in a Dataflow pipeline before importing the data to AutoML Tables.
C.Specify a weight column with higher weights for positive examples in the dataset.
D.Create duplicate copies of the positive class rows to balance the dataset.
AnswerC

A weight column lets AutoML Tables apply per-row loss multipliers, so positive examples at 5% prevalence can contribute proportionally more during training. This directly addresses the stem's imbalance constraint without resampling, preserving all 95% negative rows while preventing the model from defaulting to the majority class.

Why this answer

AutoML Tables supports a weight column that lets you assign higher importance to specific rows during training. By giving positive examples (the 5% minority class) higher weights, the model's loss function penalizes misclassification of the minority class more heavily, effectively rebalancing the learning signal without altering the dataset. This is the native, supported mechanism in AutoML Tables for handling class imbalance.

Exam trap

PDE often tests whether candidates know that AutoML Tables has a built-in weight column for class imbalance, rather than assuming external techniques like SMOTE or duplication are required.

How to eliminate wrong answers

Option A is wrong because increasing the training budget only extends training time and does not change how the model weights errors across classes, so the imbalance remains unaddressed. Option B is wrong because SMOTE must be applied outside AutoML Tables (e.g., in a Dataflow pipeline), but AutoML Tables does not natively integrate SMOTE and the recommended approach is to use the built-in weight column rather than synthetic oversampling. Option D is wrong because duplicating positive rows is a manual oversampling technique that inflates the dataset, risks overfitting to duplicated examples, and is unnecessary when the weight column provides the same effect more cleanly.

515
MCQeasy

A team has multiple versions of a model and wants to manage them centrally, including tracking metadata and promoting versions to production. Which tool should they use?

A.Cloud Storage
B.BigQuery
C.GitHub
D.Vertex AI Model Registry
AnswerD

Vertex AI Model Registry provides a centralised repository for managing model versions, tracking metadata, and controlling deployment stages, directly satisfying the stem's requirement to promote versions to production. It supports version aliases and lineage tracking, unlike raw Cloud Storage artefacts or standalone training pipelines.

Why this answer

Vertex AI Model Registry is the correct tool because it is purpose-built for centrally managing multiple model versions, tracking metadata (such as training parameters, evaluation metrics, and lineage), and promoting versions through stages like staging to production. Unlike generic storage or version control systems, it provides native integration with Vertex AI Pipelines and endpoints for controlled rollout and rollback.

Exam trap

Google often tests the misconception that a general-purpose version control system like GitHub is sufficient for ML model management, but the exam expects candidates to recognize that model registries provide specialized metadata tracking and lifecycle promotion features absent in code-only repositories.

How to eliminate wrong answers

Option A is wrong because Cloud Storage is an object storage service for raw data and artifacts, not a model management system; it lacks built-in version tracking, metadata indexing, and promotion workflows for ML models. Option B is wrong because BigQuery is a serverless data warehouse for analytics and SQL queries, not designed to store or manage ML model versions or their lifecycle. Option C is wrong because GitHub is a code repository for version control of source code and configuration files, but it does not natively handle ML model artifacts, track model-specific metadata (e.g., evaluation metrics, training hyperparameters), or provide staging-to-production promotion workflows without extensive custom tooling.

516
MCQeasy

A data engineer needs to inspect a BigQuery table for sensitive data such as credit card numbers and email addresses before sharing it with a third party. The engineer also wants to de-identify the data by masking the sensitive columns. Which Google Cloud service should be used?

A.Dataplex
B.BigQuery column-level security
C.Data Catalog
D.Cloud DLP
AnswerD

Cloud DLP inspects BigQuery data for infoTypes such as credit card numbers and email addresses, then applies de-identification transforms like masking or tokenisation. It satisfies both requirements in one service, discovery and masking, before the table is shared with the third party.

Why this answer

Cloud DLP (Data Loss Prevention) is the correct service because it is specifically designed to inspect, classify, and de-identify sensitive data such as credit card numbers and email addresses. It provides built-in infoType detectors for over 150 types of sensitive data and supports masking, tokenization, and other de-identification techniques. The engineer can use Cloud DLP to scan BigQuery tables and then apply transformations to mask the sensitive columns before sharing.

Exam trap

The trap here is that candidates confuse BigQuery column-level security (access control) with data masking, but BigQuery column-level security only hides data from unauthorized users and does not inspect or transform the data itself, whereas Cloud DLP performs actual de-identification.

How to eliminate wrong answers

Option A is wrong because Dataplex is a data fabric service for managing, governing, and cataloging data across lakes, warehouses, and marts, but it does not natively inspect or de-identify sensitive data; it relies on integration with Cloud DLP for such tasks. Option B is wrong because BigQuery column-level security (using policy tags) only controls access at the column level by granting or denying read permissions, but it does not inspect content for sensitive data or perform masking/de-identification. Option C is wrong because Data Catalog is a metadata management service for discovering and tagging data assets, but it cannot scan for sensitive data patterns or apply de-identification transformations.

517
MCQmedium

A data engineer manages a Cloud SQL for MySQL instance that stores order records. Compliance requires that the data be recoverable to any point in time within the last 30 days, and the team wants the smallest possible recovery window. They also want to avoid managing their own backup rotation scripts. Which configuration should the engineer implement?

A.Configure a read replica in a second zone and promote it if the primary fails.
B.Enable point-in-time recovery (PITR) with a transaction log retention of 30 days alongside automated backups.
C.Enable automated backups with a 30-day retention period and rely on the daily backup only.
D.Export nightly mysqldump files to a Cloud Storage bucket with a 30-day lifecycle rule.
AnswerB

Cloud SQL point-in-time recovery replays binary logs on top of the most recent automated backup, letting you restore to a specific timestamp. Setting the transaction log retention to 30 days satisfies the compliance window, and the feature is managed by the platform, so no custom rotation scripts are needed. This directly meets both the recovery granularity and the operational simplicity requirements.

Why this answer

Point-in-time recovery in Cloud SQL combines the latest automated backup with continuously archived binary logs, enabling restoration to a chosen timestamp. Extending transaction log retention to 30 days matches the compliance window while keeping the process fully managed. Daily backups alone or nightly dumps cannot achieve arbitrary-timestamp recovery, and read replicas address availability rather than historical restore.

Exam trap

The trap here is assuming that a 30-day backup retention period alone delivers 30 days of point-in-time recovery, when log retention is the separate control that governs the recovery window.

518
Multi-Selecteasy

Which TWO actions can help reduce prediction latency for a Vertex AI endpoint?

Select 2 answers
A.Increase the number of features
B.Optimize the model architecture to reduce size
C.Use a custom prediction container with optimized dependencies
D.Use a larger machine type with more vCPUs
E.Set min replicas to 0 to save cost
AnswersB, C

Smaller models predict faster.

Why this answer

Optimizing the model architecture to reduce size directly decreases the computational load during inference, which lowers prediction latency. Smaller models require fewer floating-point operations (FLOPs) per prediction, enabling faster response times on Vertex AI endpoints.

Exam trap

Google Cloud often tests the misconception that adding more compute resources (larger machine types) always reduces latency, when in fact it can increase overhead and does not address the root cause of slow inference, which is model complexity.

519
MCQmedium

Your team is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs MapReduce jobs that read and write data to HDFS. You want to minimize code changes and operational overhead. Which Google Cloud service should you use to run these jobs?

A.BigQuery
B.Cloud Composer
C.Cloud Dataproc
D.Cloud Dataflow
AnswerC

Cloud Dataproc is a fully managed service for running Apache Hadoop and Spark clusters. It supports running existing MapReduce jobs with minimal changes, as it provides HDFS-compatible storage via the cluster's local HDFS or Cloud Storage connector. Dataproc handles cluster provisioning, scaling, and management, reducing operational overhead while allowing you to lift and shift your Hadoop workloads.

Why this answer

Cloud Dataproc is designed for running Hadoop and Spark workloads in Google Cloud. It provides a managed cluster with Hadoop, Spark, and Hive, and can use Cloud Storage as a replacement for HDFS via the Cloud Storage connector. Existing MapReduce jobs can often run with little modification, and Dataproc handles cluster management, making it the best choice for migrating on-premises Hadoop workloads with minimal code changes and operational overhead.

Exam trap

The trap here is assuming that any data processing service can run Hadoop jobs, or that Dataflow can execute MapReduce code without rewriting.

520
MCQmedium

A company uses Cloud Storage as a data lake with raw, curated, and processed zones. Data in the raw zone should be automatically moved to a cheaper storage class after 30 days, and deleted after 1 year. What is the most efficient way to implement this?

A.Use Object Lifecycle Management with rules to transition to Coldline after 30 days and delete after 365 days.
B.Write a Cloud Function that runs daily, checks object ages, and moves/deletes them.
C.Use Cloud Scheduler to run a script that changes storage class and deletes objects.
D.Set a retention policy on the raw zone to prevent deletion and manually clean up.
AnswerA

Object Lifecycle Management applies rule-based transitions and deletions natively at the bucket level, so no external scheduler or code is needed. Setting Age conditions of 30 days for Coldline transition and 365 days for deletion satisfies both stem constraints automatically and efficiently.

Why this answer

Object Lifecycle Management in Cloud Storage allows you to set rules based on object age. You can transition objects to a lower-cost storage class (e.g., Nearline or Coldline) after 30 days and delete after 365 days.

521
Multi-Selectmedium

A data engineer is building a Dataflow pipeline that reads newline-delimited JSON files from a Cloud Storage bucket and writes them to BigQuery. The files arrive continuously, some with malformed records, and the team wants the pipeline to keep running while routing bad records to a dead-letter location for later analysis. The team also wants the schema to be inferred from the files during development but fixed in production. Which two approaches should the data engineer use? (Choose two.)

Select 2 answers
A.Define an explicit TableSchema in the pipeline for production and use schema autodetect only in a development branch.
B.Enable the --enableStreamingEngine flag so that schema inference happens automatically at the service level.
C.Use BigQueryIO with withSchemaUpdateOptions set to ALLOW_FIELD_ADDITION and rely on runtime inference for every file.
D.Configure BigQueryIO with withCreateDisposition set to CREATE_NEVER and withWriteDisposition set to WRITE_TRUNCATE.
E.Use a DoFn with a try/catch that emits failed parses to a side output tagged as dead-letter data.
AnswersA, E

Fixing the schema explicitly in production makes writes deterministic and prevents an unexpected field in one file from altering the table layout, while schema autodetect remains useful in development for discovering the shape of new data sources. This directly satisfies the requirement that inference be used only during development while production uses a fixed schema.

Why this answer

Dead-letter routing is implemented with a side output from a DoFn that catches parse failures, keeping the main pipeline alive. Schema stability in production is achieved by declaring an explicit TableSchema rather than letting autodetect run continuously, with autodetect reserved for development. Together these two choices satisfy both the resilience and schema-governance requirements without changing write dispositions or relying on service-level flags.

Exam trap

The trap here is conflating schema inference with pipeline execution settings, when inference is a development-time client concern and dead-lettering is a transform-level pattern.

522
MCQeasy

The exhibit shows an IAM policy for a BigQuery dataset. A Dataflow job is failing with 'Access Denied: Table ... User does not have bigquery.tables.get permission'. Which additional role should be granted to the service account?

A.roles/bigquery.admin
B.roles/bigquery.user
C.roles/bigquery.jobUser
D.roles/bigquery.dataEditor
AnswerD

Includes bigquery.tables.get.

Why this answer

The error indicates the service account lacks the `bigquery.tables.get` permission, which is required to read table metadata. `roles/bigquery.dataEditor` includes this permission along with `bigquery.tables.get`, `bigquery.tables.update`, and `bigquery.tables.export`, making it the minimal role that resolves the access denied error for a Dataflow job reading from a BigQuery table.

Exam trap

Google Cloud often tests the misconception that `roles/bigquery.user` or `roles/bigquery.jobUser` provide sufficient read access for Dataflow jobs, when in fact they lack the specific `bigquery.tables.get` permission needed for table metadata retrieval.

How to eliminate wrong answers

Option A is wrong because `roles/bigquery.admin` grants full control over BigQuery resources, including dataset deletion and IAM policy management, which is excessive and violates the principle of least privilege for a Dataflow job that only needs to read table data. Option B is wrong because `roles/bigquery.user` provides `bigquery.datasets.get` and `bigquery.jobs.create` but does not include `bigquery.tables.get`, so it would not resolve the specific permission error. Option C is wrong because `roles/bigquery.jobUser` only allows creating and managing jobs (e.g., queries) but does not grant any direct table read permissions like `bigquery.tables.get`.

523
MCQmedium

A company is building a real-time streaming pipeline to ingest clickstream events from web servers, enrich them with user profile data from Cloud Bigtable, and aggregate metrics into BigQuery. The expected throughput is 10,000 events per second with occasional spikes up to 50,000. The data must be processed with low latency (seconds) and exactly-once semantics. Which Google Cloud service should be the core processing engine?

A.Cloud Dataflow (Apache Beam runner)
B.Cloud Pub/Sub with Cloud Functions
C.Cloud Dataproc with Apache Spark Streaming
D.Cloud Data Fusion
AnswerA

Cloud Dataflow satisfies the low-latency, exactly-once requirement through its Apache Beam runner, which provides native exactly-once processing via streaming state and checkpointing. It autoscales horizontally to absorb the 10,000–50,000 events-per-second spikes, and reads Cloud Bigtable for enrichment before writing aggregates into BigQuery.

Why this answer

Cloud Dataflow, as a managed Apache Beam runner, is the correct choice because it provides exactly-once processing semantics, low-latency streaming (sub-second to seconds), and autoscaling to handle throughput spikes from 10,000 to 50,000 events per second. Its unified batch and streaming model allows you to enrich clickstream events with user profile data from Cloud Bigtable via side inputs or asynchronous lookups, and write aggregated metrics to BigQuery with exactly-once guarantees using the Beam BigQuery I/O connector.

Exam trap

Google Cloud often tests the misconception that Cloud Pub/Sub with Cloud Functions is sufficient for low-latency streaming, but candidates overlook that Cloud Functions lacks stateful processing and exactly-once semantics, making it unsuitable for aggregation and enrichment at high throughput.

How to eliminate wrong answers

Option B (Cloud Pub/Sub with Cloud Functions) is wrong because Cloud Functions has a maximum timeout of 9 minutes and does not support exactly-once processing semantics; it is at-least-once by default and lacks checkpointing for stateful operations like aggregation. Option C (Cloud Dataproc with Apache Spark Streaming) is wrong because Spark Streaming's micro-batch architecture introduces a minimum latency of several seconds (typically 5-10 seconds), which does not meet the 'seconds' low-latency requirement, and managing exactly-once semantics requires additional configuration (e.g., Kafka offsets) that is not natively handled by the managed service. Option D (Cloud Data Fusion) is wrong because it is a visual ETL tool designed for batch-oriented data integration and does not support real-time streaming ingestion or exactly-once processing; its pipelines are not suitable for sub-second latency or high-throughput event streams.

524
MCQhard

A company uses Cloud Dataproc to run Spark jobs on ephemeral clusters. The input data is in Cloud Storage and output is also to Cloud Storage. The cluster is created and deleted daily. The cost is high due to spinning up nodes. Which change can reduce cost without sacrificing performance?

A.Use standard VMs with a larger number of smaller machines
B.Use Cloud Dataflow instead
C.Use a combination of standard and preemptible VMs for worker nodes
D.Use preemptible VMs for all nodes
AnswerC

Preemptible VMs for workers reduce cost significantly; standard VMs for the master and a few worker nodes ensure reliability.

Why this answer

Using a combination of standard and preemptible VMs for worker nodes reduces cost significantly while maintaining performance. Preemptible VMs are up to 80% cheaper than standard VMs, and since Spark is fault-tolerant and can handle node preemptions via speculative execution, the job can complete without performance degradation. Standard VMs for master nodes ensure cluster stability, while preemptible workers handle the bulk of data processing.

Exam trap

Google Cloud often tests the misconception that preemptible VMs can be used for all nodes, but the trap here is that the master node must be a standard VM to avoid cluster instability, while workers can safely use preemptible VMs due to Spark's fault tolerance.

How to eliminate wrong answers

Option A is wrong because using a larger number of smaller machines increases overhead from inter-node communication and task scheduling, potentially degrading performance and not necessarily reducing cost. Option B is wrong because Cloud Dataflow is a different service for batch and stream processing, not a direct replacement for Spark on Dataproc; migrating would require rewriting jobs and may not preserve existing Spark-specific logic or performance characteristics. Option D is wrong because using preemptible VMs for all nodes, including the master node, risks cluster failure if the master is preempted, as Dataproc does not automatically recover the master; this sacrifices reliability and can cause job failures.

525
MCQhard

A team needs to run hybrid transactional/analytical workloads on PostgreSQL-compatible data with low latency. They require high performance on both OLTP and OLAP queries, leveraging a columnar engine. Which Google Cloud service is best suited?

A.AlloyDB
B.Cloud SQL for PostgreSQL
C.Cloud Spanner
D.BigQuery
AnswerA

AlloyDB satisfies the hybrid transactional/analytical requirement through its columnar engine, which accelerates OLAP scans while retaining row-based OLTP performance on PostgreSQL-compatible data. This delivers the low-latency analytics alongside transactional throughput that the team needs, unlike Cloud SQL or standard PostgreSQL offerings lacking an integrated columnar accelerator.

Why this answer

AlloyDB is the correct choice because it is a fully managed PostgreSQL-compatible database service on Google Cloud that combines a columnar engine for fast analytical queries with high transactional performance. It uses a columnar query accelerator to offload analytical workloads from the transactional engine, enabling low-latency hybrid transactional/analytical processing (HTAP) without data movement.

Exam trap

The trap here is that candidates may confuse BigQuery's columnar storage with PostgreSQL compatibility, or assume Cloud SQL's PostgreSQL support is sufficient for HTAP workloads, overlooking the need for a dedicated columnar engine.

How to eliminate wrong answers

Option B (Cloud SQL for PostgreSQL) is wrong because it lacks a columnar engine and is optimized primarily for OLTP workloads, so analytical queries would suffer from high latency and poor performance. Option C (Cloud Spanner) is wrong because it is a globally distributed, strongly consistent relational database designed for horizontal scalability and high availability, not for columnar analytics or PostgreSQL compatibility. Option D (BigQuery) is wrong because it is a serverless data warehouse with a columnar storage engine but is not PostgreSQL-compatible and is designed for OLAP, not low-latency OLTP transactions.

Page 6

Page 7 of 10

Page 8

All pages