Courseiva

Google Professional Data Engineer (PDE) — Questions 601–675

747 questions total · 10pages · All types, answers revealed

Page 8

Page 9 of 10

Page 10
601
MCQeasy

You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?

A.Cloud Dataflow with Beam SQL
B.Cloud Dataproc with Dataproc on GKE
C.Cloud BigQuery with external tables on Cloud Storage
D.Cloud Dataproc with Cloud Storage and Dataproc Metastore
AnswerD

Dataproc Metastore provides a managed Hive metastore service compatible with the Hive Metastore API, so existing Hive queries run unchanged and share metadata across multiple Dataproc clusters. Cloud Storage replaces HDFS as the storage layer, satisfying the requirement to minimise code changes during migration.

Why this answer

Cloud Dataproc runs managed Spark and Hive clusters, so existing Spark ETL jobs and Hive queries migrate with minimal code changes. Pairing Dataproc with Cloud Storage for data and Dataproc Metastore for a shared Hive metastore across clusters preserves the same table definitions and schema across multiple clusters.

Exam trap

The trap is picking a modern serverless option (Dataflow or BigQuery) that requires code rewrites, when the requirement explicitly says 'minimize changes' and 'same metastore across multiple clusters' — only Dataproc plus Dataproc Metastore satisfies both.

How to eliminate wrong answers

Option A is wrong because Dataflow with Beam SQL requires rewriting Spark/Hive logic into Beam pipelines and does not provide a Hive metastore. Option B is wrong because Dataproc on GKE changes the runtime model and does not by itself provide a shared Hive metastore across clusters. Option C is wrong because BigQuery with external tables replaces Hive semantics and requires rewriting queries, and it does not preserve a Hive metastore.

602
MCQhard

You are using Cloud Dataflow to stream data from Pub/Sub into BigQuery. The incoming messages are JSON strings with a field `event_time` in ISO 8601 format (e.g., "2024-03-15T14:30:00Z"). You need to write the data to a BigQuery table with a column `event_timestamp` of type TIMESTAMP. Which transformation should you apply in your Dataflow pipeline?

A.Use a SQL query in BigQuery to convert the string to TIMESTAMP after the data is loaded into a staging table.
B.Use a ParDo to parse the JSON and convert `event_time` to a BigQuery TIMESTAMP using the appropriate library for your pipeline language (e.g., Instant.parse in Java or datetime.fromisoformat in Python).
C.Configure BigQueryIO to write the `event_time` field as a STRING and rely on BigQuery's automatic schema detection to convert it to TIMESTAMP.
D.In the Dataflow pipeline, use a WithTimestamps transform to assign the event time as the element timestamp, and then write the original string to BigQuery.
AnswerB

Parsing the JSON and converting the ISO 8601 string to a native timestamp type ensures that BigQuery receives a proper TIMESTAMP value. Dataflow's BigQueryIO can then write it correctly. This approach handles time zones and formatting consistently, and is the standard way to transform string timestamps in a pipeline.

Why this answer

The correct approach is to parse the JSON and convert the ISO 8601 string to a native timestamp type within the Dataflow pipeline. This ensures that the data written to BigQuery matches the TIMESTAMP column type. BigQueryIO expects the input PCollection to contain TableRow objects with fields matching the destination schema, so converting the string to a timestamp before writing is necessary.

Exam trap

The trap here is assuming that BigQuery can automatically convert string timestamps during streaming inserts, which it cannot without explicit transformation.

603
MCQmedium

A data engineering team uses Cloud Data Fusion to build ETL pipelines. They have a pipeline that reads from Cloud SQL, transforms data using Wrangler, and writes to BigQuery. The pipeline fails intermittently with a 'connection timeout' error from Cloud SQL. What is the best way to handle this?

A.Use Cloud NAT to provide a static IP for Data Fusion to whitelist.
B.Configure the Cloud SQL connector in Data Fusion to use retry logic and increase the connection timeout.
C.Increase the number of Data Fusion nodes to distribute the load.
D.Migrate Cloud SQL to Cloud Spanner to handle higher concurrency.
AnswerB

Intermittent connection timeouts are transient network or database availability faults, not pipeline logic errors. Configuring the Cloud SQL connector with retry logic and a longer connection timeout lets Data Fusion reattempt failed connections automatically, resolving the intermittent failure without redesigning the pipeline.

Why this answer

Cloud Data Fusion's Cloud SQL connector can be configured with retry logic and an increased connection timeout to handle transient network issues. This directly addresses the intermittent 'connection timeout' error without requiring architectural changes, as the error is likely due to brief network latency or resource contention, not a persistent connectivity problem.

Exam trap

The trap here is that candidates often assume connectivity issues require network-level fixes (like static IPs or NAT) or scaling, rather than recognizing that transient timeouts are best handled by application-level retry and timeout configuration.

How to eliminate wrong answers

Option A is wrong because using Cloud NAT to provide a static IP for whitelisting addresses IP-based access control, but the error is a connection timeout, not an authorization failure; whitelisting does not resolve transient network delays. Option C is wrong because increasing the number of Data Fusion nodes distributes compute load but does not fix connection timeouts to Cloud SQL, which are caused by network or database-side issues, not pipeline parallelism. Option D is wrong because migrating to Cloud Spanner is an overengineered solution for a transient timeout; it introduces unnecessary complexity and cost, and does not address the root cause of intermittent connectivity.

604
MCQhard

A company uses Dataplex to manage data lakes on Google Cloud. They want to enforce data quality rules on a BigQuery table, such as ensuring that a 'email' column is not null and matches a regex pattern. Which Dataplex feature should they use?

A.Dataplex Universal Catalog
B.Dataplex Lake
C.Dataplex Data Quality
D.Dataplex Data Lineage
AnswerC

Dataplex Data Quality enforces row-level and column-level rules directly on BigQuery tables, supporting not-null checks and regex pattern matching through its built-in rule types. This satisfies the stem's requirement to validate the 'email' column's nullability and format, with results surfaced as data quality scores and scan reports.

Why this answer

Dataplex Data Quality is the feature specifically designed to define and enforce data quality rules on BigQuery tables, including checks for null values and regex pattern matching. It allows you to create data quality rules that can be scheduled or run on-demand, and it provides results that can be used for monitoring and alerting. The other options are Dataplex components for metadata management, lake organization, and lineage tracking, not for enforcing data quality rules.

Exam trap

PDE often tests the distinction between Dataplex features, and candidates may confuse Data Quality with Data Lineage or Universal Catalog, assuming that metadata management includes quality checks.

How to eliminate wrong answers

Option A is wrong because Dataplex Universal Catalog is a metadata management service that provides a unified view of data assets, not a tool for defining or enforcing data quality rules. Option B is wrong because Dataplex Lake is a logical container for organizing data assets across storage and analytics services, but it does not itself enforce data quality rules. Option D is wrong because Dataplex Data Lineage tracks data movement and transformation across systems, providing visibility into data provenance, but it does not validate data quality.

605
MCQhard

A data engineer manages a Cloud Composer 2 environment. A DAG that downloads a large reference dataset each night occasionally exceeds the default task timeout because the source API is slow. The engineer wants the task to fail fast and be retried automatically rather than hanging for hours, and wants failed runs to be visible for alerting. Which configuration should be applied to the task?

A.Configure the worker's celery_worker_autoscale setting to add more workers so the slow API call finishes sooner.
B.Set depends_on_past to true and add a retry_exponential_backoff so the task waits for the prior run before starting.
C.Increase the DAG's schedule interval and add a max_active_runs limit of one so overlapping runs cannot occur.
D.Set execution_timeout on the task to a bounded value and set retries with a retry_delay so failed attempts are re-queued and surfaced in task state.
AnswerD

execution_timeout caps how long a task may run before Airflow marks it failed, and the retries plus retry_delay settings cause the scheduler to re-queue it automatically. The failure state is recorded and available for alerting, which matches the requirement to fail fast, retry, and remain visible.

Why this answer

execution_timeout is the task-level setting that bounds wall-clock runtime and triggers a failure, while retries and retry_delay make the scheduler automatically re-queue the task. Together they produce fast failure, automatic retry, and a recorded state that alerting can observe, which is precisely what the scenario requires for a slow external API.

Exam trap

The trap here is confusing concurrency controls such as max_active_runs or worker autoscaling with a per-task execution timeout.

606
MCQmedium

A retail company is building a recommendation engine that requires processing customer clickstream data in near real-time. The data is ingested via Pub/Sub, and must be joined with a lookup table of product details (updated daily) before being used for model inference. Which design pattern should they use?

A.Enrich the stream by querying BigQuery for each event using a Cloud Function.
B.Use a Dataflow pipeline that reads from Pub/Sub and uses a side input from a regularly refreshed PCollection loaded from Cloud Storage.
C.Store product details in Cloud Memorystore (Redis) and have the streaming application look up each event.
D.Write events to BigQuery and use scheduled queries to join with the product table in batch.
AnswerB

A side input broadcasts the product lookup PCollection to every worker, letting each clickstream element be enriched in-flight without a per-element external lookup. Refreshing it daily from Cloud Storage matches the stem's daily update cadence while preserving near-real-time inference.

Why this answer

Dataflow can read streaming data from Pub/Sub and use a side input from a regularly refreshed PCollection loaded from Cloud Storage. This pattern allows the product lookup table (updated daily) to be periodically reloaded into the pipeline as a side input, enabling efficient, low-latency enrichment of each event without per-event external calls or batch delays.

Exam trap

Google Cloud often tests the distinction between streaming enrichment patterns that require external lookups (which add latency and cost) versus using side inputs for static or slowly-changing reference data, leading candidates to mistakenly choose a cache-based solution like Redis when the data is already available in Cloud Storage.

How to eliminate wrong answers

Option A is wrong because querying BigQuery for each event via a Cloud Function would introduce high latency and cost due to per-event query overhead, and BigQuery is not designed for real-time point lookups. Option C is wrong because while Cloud Memorystore (Redis) provides low-latency lookups, it requires managing a separate cache and does not natively integrate with the daily-updated Cloud Storage file; the pattern also lacks the automatic refresh mechanism that side inputs provide. Option D is wrong because writing events to BigQuery and using scheduled queries for batch joins introduces significant latency (minutes to hours), which violates the near real-time requirement for the recommendation engine.

607
MCQhard

A data scientist wants to import a pre-trained TensorFlow model into BigQuery ML for batch predictions. The model is stored in a Cloud Storage bucket. Which statement is correct?

A.Use CREATE MODEL with model_type='tensorflow' and model_path='gs://bucket/model'.
B.Use CREATE MODEL with model_type='imported_tensorflow' and model_path='gs://bucket/model'.
C.First upload the model to Vertex AI Model Registry, then reference it in BigQuery ML.
D.Use the ML.IMPORT_MODEL function to load the model into BigQuery.
AnswerA

`CREATE MODEL` with `model_type='tensorflow'` and a `gs://` `model_path` is the supported import path for TensorFlow SavedModels into BigQuery ML, satisfying the stem's requirement to load a pre-trained model from Cloud Storage for batch prediction via `ML.PREDICT`.

Why this answer

BigQuery ML supports importing TensorFlow models via CREATE MODEL with model_type='tensorflow' and a model_path pointing to a Cloud Storage location containing the SavedModel. This allows batch prediction using ML.PREDICT directly in BigQuery without moving data to Vertex AI.

Exam trap

The trap is inventing plausible-sounding model_type strings like 'imported_tensorflow' or fake functions like ML.IMPORT_MODEL — candidates who haven't memorized the exact DDL syntax fall for these.

How to eliminate wrong answers

Option B is wrong because 'imported_tensorflow' is not a valid model_type value — the correct type string is 'tensorflow'. Option C is wrong because BigQuery ML can import TensorFlow models directly from GCS; routing through Vertex AI Model Registry is unnecessary and not how BQML imports work. Option D is wrong because ML.IMPORT_MODEL is not a real BigQuery ML function — model creation uses the CREATE MODEL DDL statement.

608
MCQhard

You are designing a system to serve predictions from a large language model (LLM) with a latency SLO of 500ms. The model does not fit on a single GPU and requires model parallelism. You are considering using Vertex AI Endpoints with a custom container. What additional setup is required to achieve the latency target?

A.Compile the model using TensorFlow XLA to optimize for single GPU execution.
B.Deploy the model across multiple endpoints and use a load balancer to send requests to different parts of the model.
C.Use Vertex AI Prediction as a service for LLMs, which automatically handles hardware selection.
D.Use a machine type with multiple GPUs and configure the container to use tensor parallelism.
AnswerD

Tensor parallelism shards each model layer's weights across multiple GPUs, enabling the oversized LLM to fit and compute in parallel. A multi-GPU machine type with the container configured for tensor parallelism meets the 500ms latency SLO.

Why this answer

The model does not fit on a single GPU and requires model parallelism. Using a machine type with multiple GPUs and configuring the container to use tensor parallelism allows the model to be split across GPUs within a single instance, enabling efficient parallel computation to meet the 500ms latency SLO. Tensor parallelism distributes individual tensor operations across GPUs, reducing communication overhead compared to pipeline parallelism and is a standard approach for large models on multi-GPU instances in Vertex AI.

Exam trap

Candidates often incorrectly assume that distributing inference across multiple Vertex AI endpoints with a load balancer achieves model parallelism. However, this approach introduces network latency and cannot match the low-latency inter-GPU communication required for tensor parallelism within a single multi-GPU instance.

How to eliminate wrong answers

Option A is wrong because TensorFlow XLA compiles and optimizes computation graphs for single-device execution, but the model does not fit on a single GPU, so XLA cannot solve the memory constraint or enable model parallelism across multiple GPUs. Option B is wrong because deploying the model across multiple endpoints and using a load balancer would require splitting the model into separate services, introducing significant network latency between endpoints and breaking the model's internal state, making it impractical for the tight 500ms SLO. Option C is wrong because Vertex AI Prediction as a service for LLMs does not automatically handle custom model parallelism configurations; it provides pre-built endpoints for specific model architectures but does not support arbitrary custom containers that require tensor parallelism setup.

609
MCQeasy

You need to design a data processing system that ingests streaming data from thousands of IoT devices. The data must be processed in real-time to calculate average temperature per device over 1-minute intervals, and the results should be stored in BigQuery for analysis. You want a serverless solution with minimal management. Which combination of Google Cloud services should you use?

A.Cloud Pub/Sub for ingestion, Cloud Functions for processing, and BigQuery for storage.
B.Cloud Pub/Sub for ingestion, Cloud Dataflow for processing, and BigQuery for storage.
C.Cloud Pub/Sub for ingestion, Cloud Dataproc for processing, and Cloud Bigtable for storage.
D.Cloud IoT Core for ingestion, Cloud Dataproc for processing, and BigQuery for storage.
AnswerB

Pub/Sub is a scalable, serverless messaging service for ingesting streaming data. Dataflow is a serverless, fully managed service for stream processing that can compute 1-minute averages using windowing. BigQuery is a serverless data warehouse for storing and analyzing results. This combination requires minimal management and is ideal for real-time IoT analytics.

Why this answer

The scenario requires serverless ingestion, real-time processing with windowing, and analytical storage. Pub/Sub handles ingestion at scale, Dataflow provides serverless stream processing with windowing to compute 1-minute averages, and BigQuery stores results for analysis. This combination is fully managed and requires minimal operational effort.

Exam trap

The trap here is selecting services that are not serverless (like Dataproc) or not suited for stream processing (like Cloud Functions), or using deprecated services like Cloud IoT Core.

610
MCQmedium

A data engineer manages a BigQuery dataset holding a 40 TB partitioned table of point-of-sale transactions. Analysts frequently run queries that filter on a region column and a sale_date column together, but each query scans the full table because the region predicate is not reducing bytes billed. The engineer wants to reduce bytes scanned without changing the analytical queries or the write pipeline. What should the engineer do?

A.Define a materialized view that aggregates the table by region and sale_date and let the optimizer rewrite all analyst queries to use it.
B.Enable the require_partition_filter option on the table so that every query must include a sale_date predicate.
C.Convert the table to an external table over Cloud Storage so that only the matching Parquet row groups are read at query time.
D.Create a clustered table on the region column, because clustering co-locates rows with similar values inside each partition so block pruning can skip irrelevant data.
AnswerD

Clustering on region sorts data within each partition by that column, so BigQuery can prune storage blocks when a region predicate is present. Because the table is already partitioned on sale_date, adding clustering on region shrinks bytes scanned for queries that filter on both columns, and the write pipeline keeps inserting data unchanged.

Why this answer

Clustering is the right lever because it physically sorts data within each partition by the region column, enabling block pruning when region predicates are present. Partitioning already handles sale_date, so adding clustering on the other frequently filtered column reduces scanned bytes while leaving the queries and ingestion pipeline untouched.

Exam trap

The trap here is assuming that partitioning alone always prunes scans, when partitioning only helps for predicates on the partition column.

611
Multi-Selectmedium

A company uses Cloud Pub/Sub for event ingestion. They want to ensure that if a subscriber fails to process a message after 5 attempts, the message is sent to a dead letter topic for analysis. Which TWO configurations are needed?

Select 2 answers
A.Set max delivery attempts to 5 on the subscription.
B.Set the subscription's ack deadline to 600 seconds.
C.Enable message ordering on the subscription.
D.Create a dead letter topic and attach it to the subscription.
E.Set the subscription type to push.
AnswersA, D

Setting max delivery attempts to 5 on the subscription triggers dead-lettering once a message exceeds that threshold. Cloud Pub/Sub requires both this retry limit and a dead letter topic on the same subscription; without the limit, failed messages retry indefinitely, so this satisfies the "after 5 attempts" constraint in the stem.

Why this answer

Option A is correct because Pub/Sub's dead letter policy requires you to specify a maxDeliveryAttempts value on the subscription; setting it to 5 ensures the message is forwarded to the dead letter topic after exactly 5 failed delivery attempts. Option D is correct because a dead letter policy only takes effect when a separate dead letter topic exists and is attached to the subscription via the deadLetterPolicy.deadLetterTopic field, which is where unprocessable messages are published for analysis. Option B is not needed because the ack deadline (e.g., 600 seconds) controls how long a subscriber has to acknowledge a message before redelivery, not the number of attempts before dead-lettering.

Option C is unrelated because message ordering only preserves publish order for ordering keys and does not trigger dead-letter routing. Option E is incorrect because push versus pull is a delivery mode choice and has no bearing on the dead letter policy, which works with either subscription type.

612
MCQhard

A company wants to use Cloud DLP to inspect data in BigQuery for sensitive information and de-identify it by masking credit card numbers. They want to perform this on a schedule. Which approach should they take?

A.Use Dataplex data quality rules with a custom SQL regex
B.Use Cloud Data Loss Prevention API with Cloud Composer
C.Use BigQuery column-level security with classification
D.Use Cloud DLP inspect and de-identify jobs triggered by Cloud Scheduler
AnswerD

Cloud Scheduler triggers Cloud DLP inspect and de-identify jobs on a defined cadence, satisfying the scheduled requirement. De-identification uses the masking transformation to obscure credit card numbers while preserving data format. This combination inspects BigQuery data and applies masking without manual intervention, matching both the recurring schedule and de-identification constraints in the stem.

Why this answer

Cloud DLP can inspect BigQuery tables and de-identify using transforms like masking. Scheduling can be done via Cloud Scheduler.

613
Multi-Selectmedium

You are building a data pipeline that ingests data from on-premises into Cloud Storage, then processes it with Dataproc, and finally loads into BigQuery. You need to schedule the pipeline to run daily. The pipeline must handle occasional failures gracefully. Which THREE Google Cloud services should you use together to achieve this? (Choose 3)

Select 3 answers
A.Cloud Storage
B.Dataproc
C.Cloud Composer
D.Dataflow
E.Pub/Sub
AnswersA, B, C

Cloud Storage serves as the ingestion landing zone for on-premises data before Dataproc processing, satisfying the pipeline's first stage. It provides durable, highly available object storage that decouples ingestion from compute, letting Dataproc read source files and write results without data loss during the daily scheduled runs.

Why this answer

Cloud Storage (A) is correct because the scenario explicitly requires ingesting on-premises data into Cloud Storage as the landing/staging layer before processing. Dataproc (B) is correct because the pipeline is specified to process the data with Dataproc, Google Cloud's managed Hadoop/Spark service. Cloud Composer (C) is correct because it is the managed Apache Airflow service that provides daily scheduling, dependency orchestration, and retry/alerting logic so occasional failures are handled gracefully.

Dataflow (D) is not needed since the processing engine in this pipeline is Dataproc, not a Beam-based Dataflow job. Pub/Sub (E) is not required because the scenario describes batch ingestion into Cloud Storage rather than real-time streaming messaging.

Exam trap

PDE often tests whether candidates confuse orchestration (Composer/Airflow) with processing (Dataflow/Dataproc) — picking Dataflow because it 'processes data' ignores that the question already specifies Dataproc and asks for scheduling.

614
MCQeasy

A data engineer needs to query a BigQuery table that contains an array of structs. They want to expand the array into separate rows for each element. Which SQL function should they use?

A.STRUCT
B.UNNEST
C.ARRAY_AGG
D.SPLIT
AnswerB

UNNEST expands an array into a set of rows, one per element, which is exactly the flattening required. Placed in the FROM clause alongside the table, it correlates each struct element into its own row for querying.

Why this answer

UNNEST is the BigQuery SQL function that flattens an array into a set of rows, one per array element. When applied to an array of structs, it expands each struct into its own row, allowing the struct fields to be selected as columns.

Exam trap

The trap is confusing array construction (STRUCT, ARRAY_AGG) with array expansion (UNNEST) — candidates who see 'array of structs' sometimes reach for STRUCT or ARRAY_AGG instead of the flattening function.

How to eliminate wrong answers

Option A is wrong because STRUCT constructs a struct value; it does not expand arrays into rows. Option C is wrong because ARRAY_AGG aggregates values into an array, the opposite of what is needed. Option D is wrong because SPLIT divides a string into an array of substrings, which is unrelated to flattening arrays of structs.

615
MCQhard

You are designing a row key for Cloud Bigtable to store user activity logs. Each log entry has a timestamp (millisecond precision) and a user ID. There will be millions of writes per second from many users. To avoid hotspotting, which row key design is BEST?

A.timestamp_millis#hash(userID)
B.timestamp_millis#userID
C.userID#timestamp_millis
D.hash(userID)#userID#timestamp_millis
AnswerD

Hashing the userID spreads writes uniformly across row ranges, preventing hotspotting from sequential or timestamp-prefixed keys. Appending userID and timestamp preserves per-user chronological ordering while the hash prefix distributes millions of writes per second evenly across tablets.

Why this answer

Best because it uses a hash of the user ID as the row key prefix, which distributes writes across all Bigtable nodes and avoids hotspotting. Appending the user ID and timestamp ensures uniqueness and supports efficient queries for a specific user's logs. This design prevents the sequential timestamp from creating a single hot node, which is critical for handling millions of writes per second.

Exam trap

Google PDE often tests the misconception that placing the most selective or unique field first (like timestamp) is best for queries, but in Bigtable the row key design must prioritize write distribution over read optimization to avoid hotspotting.

How to eliminate wrong answers

Option A is wrong because placing the timestamp first causes all writes for the same millisecond to hit a single tablet server, creating a hotspot. Option B is wrong because using the raw timestamp as the prefix leads to sequential writes that overload one node, negating Bigtable's horizontal scaling. Option C is wrong because while userID as prefix distributes writes, it does not guarantee uniqueness for multiple log entries from the same user at the same millisecond, and it lacks the hash to prevent skewed access patterns if user IDs are sequential or predictable.

616
Multi-Selectmedium

A data engineer is designing a batch processing pipeline that runs daily. The pipeline reads CSV files from GCS, transforms them using Python, and writes the results to BigQuery. They need to parameterize the pipeline for different environments and run it on a schedule. Which THREE components should they use? (Choose 3)

Select 3 answers
A.Cloud Composer
B.Dataproc
C.Dataflow Flex Template
D.Storage Transfer Service
E.Cloud Functions
AnswersA, B, C

Cloud Composer orchestrates and schedules the pipeline, and allows parameterization via Airflow variables.

Why this answer

Cloud Composer (A) is correct because it is a managed Apache Airflow service that provides DAG-based scheduling and parameterization, making it ideal for orchestrating a daily batch pipeline with environment-specific variables. Dataproc (B) is correct because it offers managed Spark/Hadoop clusters that can run Python-based transformations over the CSV data read from GCS at batch scale. Dataflow Flex Template (C) is correct because it packages a Dataflow pipeline (including Python transforms) into a reusable, parameterized template that can be invoked with different runtime parameters per environment.

Storage Transfer Service (D) is not appropriate here because it only moves data between storage systems and performs no transformation or scheduling logic. Cloud Functions (E) is unsuitable because it is an event-driven, short-lived serverless function, not designed for orchestrating daily batch pipelines or large-scale data transformation.

Exam trap

Google often tests the distinction between orchestration/scheduling services (Cloud Composer) and compute/processing services (Dataproc, Cloud Functions), leading candidates to mistakenly choose Dataproc for scheduling or Cloud Functions for batch processing.

617
Matchingmedium

Match each Google Cloud monitoring/logging service to its function.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Metrics and alerting for cloud resources

Centralized log storage and analysis

Aggregates and analyzes application errors

Records administrative and data access activities

Why these pairings

Cloud Monitoring provides metrics and dashboards; Cloud Logging manages logs; Error Reporting aggregates errors; Cloud Trace provides tracing. Common mistakes include mixing up Logging with Trace and Monitoring with Error Reporting.

618
MCQmedium

A data pipeline uses Cloud Pub/Sub to ingest events and Cloud Functions to transform and write to BigQuery. The system is experiencing data loss during Pub/Sub subscription outages. Which design change improves reliability?

A.Use Dataflow with at-least-once delivery and checkpointing
B.Use a pull subscription with a custom app that polls frequently
C.Use long ack deadlines to keep messages in the subscription
D.Increase the timeout in Cloud Functions
AnswerA

Dataflow provides exactly-once semantics with checkpointing to prevent data loss.

Why this answer

Dataflow with at-least-once delivery and checkpointing ensures that messages are not lost during Pub/Sub subscription outages because Dataflow tracks processing progress via checkpoints and can replay unacknowledged messages from the last checkpoint. This decouples the processing from the subscription's transient failures, providing fault-tolerant, exactly-once or at-least-once semantics depending on the sink.

Exam trap

Google Cloud often tests the misconception that increasing timeouts or ack deadlines alone can prevent data loss, when in reality they only delay the inevitable loss without a replay mechanism like checkpointing or a persistent buffer.

How to eliminate wrong answers

Option B is wrong because a pull subscription with a custom app that polls frequently does not inherently provide reliability during subscription outages; the app would still lose messages if the subscription itself is unavailable or if the app fails to acknowledge before the ack deadline. Option C is wrong because long ack deadlines only keep messages in the subscription for a longer time, but they do not prevent data loss if the subscriber crashes or the subscription becomes unavailable; messages can still be dropped if the deadline expires without ack. Option D is wrong because increasing the timeout in Cloud Functions does not address data loss from subscription outages; it only allows the function to run longer before timing out, but does not provide replay or checkpointing mechanisms.

619
MCQmedium

An organization runs periodic Apache Spark jobs on Dataproc to process data from Cloud Storage. They want to reduce costs by using preemptible instances for worker nodes. What is a key consideration when using preemptible instances in Dataproc?

A.Preemptible instances cannot be used with the standard cluster mode
B.Jobs must be designed to handle node preemption, and overall job runtime may increase
C.Preemptible instances are only available in certain regions
D.Jobs will automatically restart from the last checkpoint without any performance impact
AnswerB

Preemptible nodes can be reclaimed at any time, so Spark jobs must tolerate losing workers mid-execution, and recomputation plus retries typically lengthen total runtime. This matches the stem's cost-reduction goal while acknowledging the resilience and duration trade-offs preemptible instances impose.

Why this answer

Preemptible VMs in Dataproc can be reclaimed by Compute Engine at any time (with a 30-second shutdown notice), so Spark jobs must be designed to tolerate worker loss — for example, by using checkpointing, retries, or resilient data pipelines. Because preempted nodes are removed and replacements take time to spin up, overall job runtime typically increases compared to a cluster of standard VMs. This trade-off is the key operational consideration when choosing preemptible workers for cost savings.

Exam trap

PDE often tests the misconception that preemptible instances are transparent to the job — candidates must recognize that preemption causes task re-execution and longer runtimes, and that checkpointing is a developer responsibility, not an automatic Dataproc feature.

How to eliminate wrong answers

Option A is wrong because preemptible instances can absolutely be used with Dataproc standard clusters — in fact, Dataproc explicitly supports mixing preemptible workers with standard workers in a standard cluster (they cannot be the sole worker type in some configurations, but they are supported). Option C is wrong because preemptible VMs are available in most Compute Engine regions and zones, not restricted to a small subset — availability is broad, though capacity can vary. Option D is wrong because jobs do not automatically restart from the last checkpoint without performance impact; checkpointing must be explicitly implemented by the developer, and even then, restarting and re-processing increases runtime and resource consumption.

620
MCQeasy

A company needs to deploy a trained model for real-time predictions with low latency. Which Vertex AI resource should they use?

A.Cloud TPU
B.Vertex AI Batch Prediction
C.Vertex AI Endpoints
D.Cloud Run
AnswerC

Endpoints host trained models for online serving, exposing a REST/gRPC interface that returns synchronous predictions per request. This satisfies the low-latency real-time constraint, unlike batch prediction, which processes asynchronous jobs over stored data and cannot serve individual requests on demand.

Why this answer

Vertex AI Endpoints are designed for online prediction, providing a managed service that hosts models for real-time inference with low latency. They automatically scale resources and handle traffic routing, making them the correct choice for deploying a trained model that needs to respond to individual prediction requests quickly.

Exam trap

Google Cloud often tests the distinction between batch and online prediction, and the trap here is that candidates confuse Vertex AI Batch Prediction (which is for offline, large-scale inference) with the real-time serving capability of Vertex AI Endpoints, leading them to select option B.

How to eliminate wrong answers

Option A is wrong because Cloud TPUs are specialized hardware accelerators for training and batch inference, not a deployment service for real-time predictions; they require manual management and are not designed for low-latency serving of individual requests. Option B is wrong because Vertex AI Batch Prediction is intended for asynchronous, high-throughput predictions on large datasets, not for real-time, low-latency responses; it processes jobs in batches and returns results to a storage location. Option D is wrong because Cloud Run is a serverless compute platform for containerized applications, but it lacks the native model hosting, versioning, and traffic splitting capabilities that Vertex AI Endpoints provide for machine learning models.

621
MCQmedium

You need to analyse streaming data from thousands of IoT devices, each sending temperature readings every second. You want to calculate the average temperature per device over the last 5 minutes, updating every minute. Which windowing strategy should you use in Dataflow?

A.Sliding windows of length 5 minutes with a period of 1 minute
B.Global windows with a trigger firing every minute
C.Fixed windows of 5 minutes
D.Session windows with a gap duration of 1 minute
AnswerA

Sliding windows of five minutes with a one-minute period recompute overlapping five-minute averages every minute, so each device's rolling average refreshes continuously. This matches the required five-minute lookback and one-minute update cadence, which fixed or session windows cannot provide.

Why this answer

Sliding windows of length 5 minutes with a period of 1 minute produce a new window every minute, each covering the previous 5 minutes of data — exactly matching the requirement to compute a 5-minute average that updates every minute. This is the canonical use case for sliding windows in Dataflow/Beam.

Exam trap

PDE often tests the confusion between sliding and fixed windows — candidates must recognize that 'updating every minute over the last 5 minutes' requires overlapping windows (sliding), not non-overlapping fixed windows or global windows with triggers.

How to eliminate wrong answers

Option B is wrong because global windows encompass the entire unbounded dataset and never close on their own; a trigger firing every minute would emit cumulative results over all data seen so far, not a 5-minute rolling average. Option C is wrong because fixed (tumbling) windows of 5 minutes are non-overlapping and emit only once every 5 minutes, so the average would not update every minute. Option D is wrong because session windows are defined by gaps in activity — with IoT devices sending every second, sessions would rarely close, and the gap duration of 1 minute does not produce a rolling 5-minute average.

622
MCQeasy

A team is deploying a model on AI Platform Prediction. They want to monitor for data drift to maintain model quality. Which service should they use?

A.Cloud DLP
B.AI Platform Continuous Evaluation
C.Cloud Monitoring
D.Cloud Audit Logs
AnswerB

AI Platform Continuous Evaluation monitors deployed models for data drift and skew by comparing serving inputs against training baselines, directly satisfying the drift-monitoring requirement. It reports divergence metrics so the team can retrain before model quality degrades.

Why this answer

AI Platform Continuous Evaluation (CE) is the correct service because it is specifically designed to monitor deployed models for data drift and feature skew. It automatically compares the distribution of incoming prediction requests against the training data distribution, alerting when statistically significant drift is detected, which directly addresses the need to maintain model quality over time.

Exam trap

Google often tests the misconception that general-purpose monitoring or logging services (like Cloud Monitoring or Audit Logs) are sufficient for ML-specific drift detection, when in fact only a dedicated ML evaluation service like AI Platform Continuous Evaluation provides the necessary statistical comparison against training data.

How to eliminate wrong answers

Option A is wrong because Cloud DLP (Data Loss Prevention) is used to inspect, classify, and de-identify sensitive data, not to monitor statistical distributions of model features. Option C is wrong because Cloud Monitoring collects metrics and logs for infrastructure and application performance but lacks built-in statistical drift detection for ML model features. Option D is wrong because Cloud Audit Logs record administrative actions and access to resources, not the distributional properties of prediction data.

623
MCQeasy

A company needs to stream data from a fleet of IoT devices to BigQuery for near-real-time analytics. The data volume is unpredictable and can spike during certain events. Which Google Cloud service should be used as the ingestion point to handle variable throughput with minimal operational overhead?

A.Cloud Datastore
B.Cloud Functions
C.Cloud Storage
D.Cloud Pub/Sub
AnswerD

Cloud Pub/Sub decouples ingestion from processing, buffering unpredictable spikes automatically without provisioning capacity, which satisfies the variable-throughput and minimal-operational-overhead constraints. Its globally distributed, serverless design scales elastically per message volume, then feeds BigQuery through Dataflow, avoiding the fixed-capacity limits of alternatives such as Pub/Sub Lite or direct streaming inserts.

Why this answer

Cloud Pub/Sub is the correct choice because it is a fully managed, scalable messaging service designed to decouple data producers from consumers, handling unpredictable and spiky throughput without requiring manual scaling. It can ingest millions of messages per second and buffer them until BigQuery is ready to consume, ensuring near-real-time analytics with minimal operational overhead.

Exam trap

Google Cloud often tests the misconception that Cloud Functions can serve as a direct ingestion point for streaming data, but candidates overlook that Cloud Functions lacks durable buffering and automatic scaling for high-throughput spikes, making Pub/Sub the correct decoupling layer.

How to eliminate wrong answers

Option A is wrong because Cloud Datastore is a NoSQL document database for storing structured data, not a streaming ingestion service; it cannot handle variable-throughput message ingestion or buffer spikes. Option B is wrong because Cloud Functions is a serverless compute platform for event-driven code execution, not a durable ingestion buffer; it lacks built-in buffering and would require custom scaling logic to handle throughput spikes. Option C is wrong because Cloud Storage is an object storage service for batch data, not designed for near-real-time streaming ingestion; it introduces latency and requires additional components (e.g., Cloud Functions or Pub/Sub notifications) to trigger downstream processing.

624
Multi-Selecthard

A media company runs a Cloud Composer environment whose DAGs trigger Dataflow batch jobs and BigQuery loads. The operations team reports that Composer costs are rising and that DAGs occasionally stall because workers are saturated. You review the environment and find that several tasks are long-running sensors that hold worker slots while waiting on external conditions. Which TWO changes should you make to reduce worker saturation and cost? (Choose two.)

Select 2 answers
A.Move the long-running Dataflow job monitoring out of the DAG by having the DAG submit the job and exit while a separate process tracks completion.
B.Change the DAG schedule from a cron expression to a timedelta interval so runs are spaced further apart.
C.Set the Composer environment's worker count to the minimum and rely on autoscaling to add workers only when the queue is deep.
D.Increase the DAG's max_active_tasks parameter so more tasks can run in parallel on the same workers.
E.Switch the long-waiting sensors to deferrable mode so they release the worker slot while waiting and resume when the condition is met.
AnswersA, E

When a task blocks a worker for the entire duration of a Dataflow job, that slot is unavailable for other work. Submitting the job and returning, then checking status separately or through a deferrable waiter, reduces the time a worker is occupied. This shortens the critical path of worker occupancy and alleviates the saturation the team is seeing.

Why this answer

Worker saturation caused by long-blocking sensors and job-waiting tasks is best relieved by changing how those tasks consume resources rather than by adding concurrency limits or rescheduling. Deferrable sensors release slots during waits, and decoupling long Dataflow job monitoring from the worker keeps the pool available for other tasks, which lowers both stalls and cost.

Exam trap

The trap here is treating worker saturation as a scheduling or concurrency problem, when the real cause is tasks that occupy worker slots for long durations while doing nothing but waiting.

625
MCQeasy

Based on the exhibit, what is the most likely cause of the out-of-memory error?

A.The BigQuery output table schema does not match the transformed data, causing write failures.
B.The Pub/Sub subscription is not acknowledging messages quickly enough, causing a backlog.
C.The worker machine type has insufficient memory for the message size and throughput.
D.The fixed window duration of 1 minute is too short, causing excessive state overhead.
AnswerC

Each Dataflow worker buffers elements in memory during processing and shuffling, so message size multiplied by throughput determines the memory footprint. If that exceeds the worker machine type's available RAM, the worker throws an out-of-memory error, making insufficient worker memory the direct cause.

Why this answer

The out-of-memory error in a Dataflow pipeline is most likely caused by the worker machine type having insufficient memory for the message size and throughput. When messages are large or the throughput is high, each worker must hold data in memory for processing, windowing, and shuffling. If the worker's memory is too small, the JVM heap runs out of memory, leading to an OOM error.

Exam trap

Google Cloud often tests the misconception that OOM errors are caused by schema mismatches or Pub/Sub backlogs, but the real cause is almost always insufficient worker memory for the data volume.

How to eliminate wrong answers

Option A is wrong because a schema mismatch between the BigQuery output table and the transformed data would cause write failures or errors in the BigQuery IO connector, not an out-of-memory error on the worker. Option B is wrong because a Pub/Sub subscription not acknowledging messages quickly enough would cause a backlog and increase unacknowledged message count, but it would not directly cause an out-of-memory error on the Dataflow worker; the pipeline would still process messages at its own pace, and the backlog would be in Pub/Sub, not in worker memory. Option D is wrong because a fixed window duration of 1 minute being too short would increase state overhead only if the pipeline uses stateful processing or triggers that accumulate state across windows; for a simple streaming pipeline, shorter windows actually reduce the amount of data held in memory per window, not cause OOM.

626
MCQhard

A financial services firm runs a batch risk calculation on Dataproc. The job reads from Cloud Storage, processes data in memory, and writes results to BigQuery. The job must complete within a 2-hour window each night, and the cluster must be shut down automatically after completion to minimize cost. You want to orchestrate this with minimal operational overhead. What should you do?

A.Create a Dataproc workflow template that runs the job and deletes the cluster upon completion, and trigger it with a Cloud Scheduler job that calls the Dataproc API.
B.Submit the job to an ephemeral Dataproc cluster created with the gcloud dataproc clusters create command, and use a Cloud Function triggered by Cloud Scheduler to delete the cluster after a fixed 2-hour delay.
C.Use a persistent Dataproc cluster and submit the job via a cron job on the master node, relying on the cluster's idle timeout to shut it down.
D.Use Cloud Composer to create a Dataproc cluster, submit the job, and delete the cluster in a DAG, scheduling it daily.
AnswerA

Dataproc workflow templates can define a job and a cluster that is automatically deleted after the workflow finishes, which meets the auto-shutdown requirement. Cloud Scheduler can invoke the Dataproc API on a cron schedule, providing orchestration with minimal operational overhead. This combination directly addresses both the scheduling and cost cleanup needs.

Why this answer

Dataproc workflow templates natively support running a job on a cluster that is deleted automatically when the workflow completes, satisfying the auto-shutdown and cost requirements. Triggering the template with Cloud Scheduler provides a simple, serverless schedule. Persistent clusters, Composer, or fixed-delay deletion add cost or operational complexity without the same guarantee of completion-based cleanup.

Exam trap

The trap here is assuming that any scheduler plus a cluster deletion script is equivalent, when only a workflow template provides native completion-based cluster deletion.

627
MCQmedium

A company stores IoT sensor data in BigQuery. Queries that filter on a timestamp column and a device_id column are slow even though the table is partitioned by day. What should the data engineer do to improve query performance?

A.Increase the partition size to monthly
B.Switch to ingestion-time partitioning instead of column-based
C.Enable automatic query rewriting with BI Engine
D.Cluster the table on device_id
AnswerD

Partitioning by day prunes on the timestamp filter, but device_id predicates still scan every partition. Clustering sorts data within each partition by device_id, so BigQuery reads only the relevant blocks, cutting bytes scanned and improving filter performance.

Why this answer

Clustering on device_id organizes the data within each day partition by device_id, allowing BigQuery to prune blocks during queries that filter on that column. This reduces the amount of data scanned and improves query performance without changing the partitioning scheme. Partitioning alone only limits scans by time range; clustering adds intra-partition sorting for non-time-based filters.

Exam trap

Google Cloud often tests the distinction between partitioning (which prunes by time) and clustering (which prunes by non-time columns), and the trap here is assuming that partitioning alone is sufficient for all filter columns, leading candidates to choose an option that changes the partition strategy rather than adding clustering.

How to eliminate wrong answers

Option A is wrong because increasing partition size to monthly would reduce the number of partitions, making each partition larger and actually increasing the data scanned for queries that filter on a specific day, worsening performance. Option B is wrong because ingestion-time partitioning is equivalent to partitioning on a pseudo-column (_PARTITIONTIME) and does not address the need to optimize filtering on device_id; it would not improve performance for queries filtering on device_id. Option C is wrong because BI Engine accelerates sub-second queries on small to medium datasets by caching results, but it does not reduce the amount of data scanned for large tables or optimize filtering on device_id; it is designed for interactive analytics, not for improving slow queries due to full table scans.

628
MCQmedium

A data scientist needs to provide explanations for each prediction made by a deployed autoML model to comply with regulatory requirements. Which Vertex AI feature should they use?

A.Vertex AI Model Monitoring
B.Vertex AI Vizier
C.Vertex AI Explainable AI
D.Vertex AI Feature Store
AnswerC

Vertex AI Explainable AI attaches feature attributions to each prediction, using methods such as Sampled Shapley for AutoML models. This satisfies the regulatory requirement for per-prediction explanations, since it returns the contribution of each input feature alongside every individual inference.

Why this answer

Vertex AI Explainable AI is the correct feature because it provides feature attributions and explanations for each prediction, enabling compliance with regulatory requirements that demand interpretability. It uses techniques like Shapley value approximations or integrated gradients to quantify the contribution of each input feature to the model's output, which is essential for auditing and transparency in deployed autoML models.

Exam trap

Google Cloud often tests the distinction between monitoring (detecting drift) and explaining (interpreting predictions), so candidates mistakenly choose Model Monitoring when the question explicitly asks for per-prediction explanations for regulatory compliance.

How to eliminate wrong answers

Option A is wrong because Vertex AI Model Monitoring is designed to detect prediction drift, data drift, and feature skew over time, not to provide per-prediction explanations. Option B is wrong because Vertex AI Vizier is a hyperparameter tuning and optimization service that helps find the best model architecture or parameters, not a tool for explaining individual predictions. Option D is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature data, but it does not generate explanations for model predictions.

629
MCQmedium

You need to automate retraining of a model when new training data becomes available every week. The training pipeline runs on Vertex AI Pipelines and is triggered by Cloud Composer. After retraining, you want to evaluate the new model against a golden dataset. If the model's accuracy improves by at least 1%, it should be automatically deployed to the staging endpoint. What is the best way to implement the decision logic?

A.Use Cloud Functions to compare metrics and call the endpoint if conditions are met.
B.Add a conditional step in the Vertex AI Pipeline to evaluate the model and deploy if the accuracy improvement threshold is met.
C.After training, run a batch prediction job on the golden dataset and compare metrics manually.
D.Use Vertex AI Experiments to log metrics and set up an alert to manually deploy.
AnswerB

A conditional step inside the Vertex AI Pipeline evaluates the freshly trained model against the golden dataset and, using pipeline conditions, deploys to the staging endpoint only when accuracy improves by at least 1%. This keeps the decision logic within the pipeline that Cloud Composer already triggers.

Why this answer

Vertex AI Pipelines supports conditional execution natively via the `Condition` component, allowing you to evaluate the new model's accuracy against the golden dataset within the same pipeline and deploy only if the improvement threshold (≥1%) is met. This approach keeps the entire retraining, evaluation, and deployment workflow automated, auditable, and tightly coupled within a single orchestrated pipeline, avoiding external triggers or manual steps.

Exam trap

Google Cloud often tests the misconception that external services like Cloud Functions are needed for decision logic, when in fact Vertex AI Pipelines' native conditional steps are the simpler, more integrated, and recommended approach for automated model evaluation and deployment within a pipeline.

How to eliminate wrong answers

Option A is wrong because Cloud Functions would introduce an external, event-driven component that adds latency, complexity, and potential failure points; Vertex AI Pipelines already provides built-in conditional logic for this exact use case, making an extra function unnecessary. Option C is wrong because running a batch prediction job and manually comparing metrics defeats the automation goal and introduces human error and delay, which is not suitable for a weekly retraining cadence. Option D is wrong because Vertex AI Experiments is designed for tracking and comparing experiments, not for automated decision-making or deployment; relying on alerts for manual deployment contradicts the requirement for automatic retraining and deployment.

630
Multi-Selectmedium

A company wants to implement model monitoring for a deployed classification model. Which three types of monitoring should they set up? (Select 3)

Select 3 answers
A.Infrastructure cost monitoring
B.Training-serving skew
C.Prediction drift
D.Input feature drift
E.Model version comparison
AnswersB, C, D

Training-serving skew monitoring compares the feature distributions your model saw during training against those arriving at prediction time, exposing preprocessing mismatches. Vertex AI Model Monitoring supports this as a distinct skew type, satisfying the requirement to detect divergence between training and serving data.

Why this answer

For a deployed classification model, the three essential monitoring types are training-serving skew (B), prediction drift (C), and input feature drift (D). Training-serving skew (B) is correct because it detects discrepancies between the feature transformations applied during training and those applied at inference time, which can silently degrade model accuracy in production. Prediction drift (C) is correct because it tracks changes in the distribution of the model's output predictions over time, signaling that the model's behavior is shifting even when ground truth labels are unavailable.

Input feature drift (D) is correct because it monitors changes in the distribution of incoming feature values relative to the training data, which is often an early warning that the model is operating on data unlike what it learned from. Infrastructure cost monitoring (A) is not a model-quality monitoring type and does not detect model degradation, and model version comparison (E) is a deployment/management activity rather than an ongoing monitoring category for a deployed classifier.

Exam trap

Candidates often confuse operational tasks (like cost monitoring or version management) with model monitoring, leading them to incorrectly select options like infrastructure cost monitoring or model version comparison.

631
Multi-Selectmedium

A data engineer is designing a Cloud Bigtable schema for high-volume time-series data. Which TWO practices should they follow to avoid performance issues?

Select 2 answers
A.Place the timestamp as the first component of the row key
B.Create as many column families as possible
C.Use a hashed prefix in the row key to distribute writes
D.Group related columns into column families
E.Store all columns in a single column family
AnswersC, D

Hashing the row key prefix scatters sequential writes across many tablets instead of concentrating them on one, preventing hot-spotting. This satisfies the high-volume time-series constraint where monotonically increasing keys would otherwise bottleneck a single tablet.

Why this answer

Option C is correct because Cloud Bigtable sorts rows lexicographically by row key, so monotonically increasing keys (like raw timestamps) concentrate writes on a single tablet, creating a hotspot; adding a hashed prefix distributes writes across multiple tablets and improves throughput. Option D is correct because grouping related columns into a small number of column families keeps data that is accessed together physically co-located, which improves read efficiency and avoids excessive per-row overhead. Option A is incorrect because placing the timestamp first produces sequential, monotonically increasing keys that cause write hotspots rather than distributing load.

Option B is incorrect because Bigtable recommends keeping the number of column families small (typically fewer than 10), since each family adds memory and compaction overhead. Option E is incorrect because cramming all columns into one family prevents efficient access patterns and mixes unrelated data, hurting read performance and locality.

Exam trap

The trap is assuming that timestamp-first row keys are good for time-series data, but they cause hotspots; candidates may also overuse column families.

632
MCQmedium

A company uses Dataproc Serverless for Spark batch jobs. They notice that some jobs are failing due to out-of-memory (OOM) errors. Which configuration parameter should they adjust to allocate more memory per executor?

A.Use a custom image with more memory
B.Set spark.driver.memory to a higher value
C.Increase the number of workers by setting --num-workers
D.Set spark.executor.memory to a higher value, e.g., 8g
AnswerD

spark.executor.memory controls the heap size allocated to each Spark executor JVM. Raising it to 8g gives executors more working memory, directly addressing the OOM failures caused by insufficient per-executor heap during Dataproc Serverless batch execution.

Why this answer

OOM errors in Spark executors are resolved by increasing executor memory, which is controlled by spark.executor.memory. Setting it to a higher value such as 8g allocates more heap to each executor, allowing it to process larger partitions without running out of memory. This is the direct and correct configuration for the reported symptom.

Exam trap

PDE often tests the confusion between driver memory and executor memory — candidates pick spark.driver.memory because it sounds like 'the memory setting' without distinguishing which process is failing.

How to eliminate wrong answers

Option A is wrong because a custom image with more memory does not change the JVM heap allocation for executors — memory must be configured via Spark properties, not the container image. Option B is wrong because spark.driver.memory affects the driver process, not executors; OOM in executors will not be fixed by increasing driver memory. Option C is wrong because increasing the number of workers adds more executors but does not increase memory per executor — if each executor still OOMs, more of them will also fail.

633
Multi-Selecthard

You are building a Dataflow pipeline that reads from Pub/Sub, applies transformations, and writes to BigQuery. The pipeline must handle late-arriving data and ensure that the windowing and triggering are correct. Which THREE configurations should you consider? (Choose 3)

Select 3 answers
A.Enable Dataflow Streaming Engine for exactly-once processing.
B.Use side inputs to enrich streaming data with static data.
C.Use the BigQuery Storage Write API with committed mode to ensure exactly-once writes.
D.Set an allowed lateness duration to handle late-arriving data.
E.Configure a triggering frequency to control how often results are emitted.
AnswersC, D, E

The BigQuery Storage Write API's committed mode provides exactly-once semantics through write-stream offsets, preventing duplicate rows when retries occur after late data triggers additional pane firings. This satisfies the pipeline's requirement for correct windowing and triggering, since speculative or repeated emissions from allowed lateness would otherwise insert duplicates into BigQuery.

Why this answer

Option C is correct because the BigQuery Storage Write API in committed mode (exactly-once semantics) uses write streams with offsets and client-side deduplication, which is the recommended way to guarantee exactly-once writes from a Dataflow streaming pipeline into BigQuery. Option D is correct because setting an allowed lateness duration (via withAllowedLateness) on the windowing strategy tells Dataflow how long to retain window state and keep accepting late-arriving elements after the watermark passes the window end, which is exactly the requirement for handling late data. Option E is correct because configuring a triggering frequency (for example, AfterProcessingTime.pastFirstElementInPane().plusDelayOf(...) or a repeated trigger) controls how often intermediate or final results are emitted from each window, which directly governs the windowing/triggering behavior described in the scenario.

Option A does not belong because Dataflow Streaming Engine is a runner feature that offloads state and windowing execution to the service for scalability and cost benefits; it does not itself provide exactly-once processing, which comes from the pipeline's sources, sinks, and deduplication logic. Option B does not belong because side inputs are used to enrich streaming records with additional (often static or slowly changing) data, which is unrelated to handling late-arriving data or to windowing and triggering correctness.

Exam trap

A common misconception is that Dataflow Streaming Engine alone provides exactly-once processing. In reality, it is the combination of source/sink semantics (like the BigQuery Storage Write API) that ensures exactly-once, not the engine itself.

634
MCQhard

A financial services firm runs an Apache Airflow workload on Cloud Composer 2 that ingests market data, runs dbt transformations, and loads curated tables into BigQuery. The DAG currently uses a single PythonOperator that runs a long shell command, and the team wants to make failures easier to diagnose and retries more granular. They also want to avoid rerunning already-successful upstream steps. Which change BEST meets these goals?

A.Move the shell command into a Cloud Function and invoke it from the PythonOperator using an HTTP call, keeping the DAG structure unchanged.
B.Split the monolithic PythonOperator into a chain of task-level operators such as BigQueryInsertJobOperator and DataprocSubmitJobOperator, with retries configured per task and depends_on_past set appropriately.
C.Increase the worker count and worker machine type on the Cloud Composer environment so the single PythonOperator task has more resources and completes faster.
D.Set the DAG's max_active_runs to 1 and enable catchup so that failed runs are automatically retried on the next schedule interval.
AnswerB

Breaking the DAG into discrete task-level operators lets Airflow retry only the failed step and preserves successful upstream results, since each task tracks its own state. Using purpose-built operators such as BigQueryInsertJobOperator gives clearer logs and status per step, and configuring retries at the task level provides granular control. This directly addresses diagnosability and avoids rerunning completed work, which is the goal described.

Why this answer

Decomposing a monolithic task into discrete operators is the standard Airflow pattern for granular retries and clearer observability. Each task maintains its own state, so a failure in the load step can be retried without rerunning ingestion or transformation. Purpose-built operators emit structured logs and status, and task-level retry configuration gives precise control over failure handling.

Exam trap

The trap here is assuming that scaling the Composer environment or enabling catchup improves retry granularity, when those settings affect capacity and scheduling rather than task-level state.

635
Multi-Selecteasy

A company wants to use BigQuery for analytics. They need to meet compliance requirements by encrypting data at rest with a key they control. Which TWO actions should they take? (Choose 2.)

Select 2 answers
A.Set the Cloud KMS key as the default encryption key for the BigQuery dataset.
B.Create a Cloud Storage bucket and load data there.
C.Use VPC Service Controls to restrict access to the dataset.
D.Create a key ring and cryptographic key in Cloud KMS.
E.Enable BigQuery column-level encryption using AEAD functions.
AnswersA, D

Setting a Cloud KMS key as the dataset default encryption key satisfies the requirement for customer-controlled keys, since BigQuery then wraps data with that CMEK rather than Google-managed encryption. This applies the key at dataset scope, so every table created within it inherits the same customer-managed protection automatically.

Why this answer

Option D is correct because customer-managed encryption keys (CMEK) in BigQuery require you to first create a Cloud KMS key ring and a cryptographic key (with appropriate rotation and IAM permissions granted to the BigQuery service account) before it can be used. Option A is correct because, once the KMS key exists, you set it as the default encryption key on the BigQuery dataset so that all tables created in that dataset are encrypted at rest with the customer-controlled key. Option B is incorrect because a Cloud Storage bucket is not required to apply CMEK to BigQuery datasets and does not by itself satisfy the BigQuery encryption requirement.

Option C is incorrect because VPC Service Controls provide network perimeter controls against data exfiltration, not encryption of data at rest. Option E is incorrect because AEAD functions provide column-level encryption of specific values within queries, not dataset-level encryption at rest with a customer-managed KMS key.

Exam trap

The trap is confusing column-level encryption (AEAD functions) with dataset-level CMEK — candidates may think encrypting columns satisfies 'encryption at rest with a key they control,' but the exam expects dataset-level CMEK configuration.

636
MCQhard

A data science team uses Vertex AI Pipelines to automate retraining. They want to ensure that only models with performance above a threshold are deployed. Which component should they add to the pipeline?

A.Vertex AI Feature Store
B.Vertex AI Model Evaluation
C.Cloud Build trigger
D.Cloud Monitoring alert
AnswerB

Vertex AI Model Evaluation computes metrics against a dataset, producing an evaluation resource. Adding it to the pipeline lets a conditional deployment step compare those metrics to the threshold, so only models meeting the performance bar are deployed.

Why this answer

Vertex AI Model Evaluation provides built-in evaluation metrics and threshold-based validation that can be used as a pipeline condition to gate model deployment. By adding a Model Evaluation component, the pipeline can compare model performance against a predefined threshold and only proceed to deploy if the metrics (e.g., AUC, precision, recall) meet or exceed the required value.

Exam trap

The trap here is that candidates may confuse monitoring (Cloud Monitoring) or feature management (Feature Store) with the evaluation step needed to gate deployment, but only Model Evaluation provides the threshold-based conditional logic within the pipeline itself.

How to eliminate wrong answers

Option A is wrong because Vertex AI Feature Store is a centralized repository for storing, serving, and sharing feature data, not for evaluating model performance or enforcing deployment thresholds. Option C is wrong because Cloud Build trigger is used to automate builds and tests of source code, not to evaluate trained model metrics within a Vertex AI Pipeline. Option D is wrong because Cloud Monitoring alert is designed to notify operators about system or application anomalies, not to serve as a pipeline gate that conditionally deploys models based on evaluation results.

637
Multi-Selectmedium

A company wants to build a reporting pipeline where data is collected from IoT devices, stored raw in Cloud Storage, and then processed into BigQuery for analytics. They need to ensure data is encrypted at rest using customer-managed keys. Which THREE steps should they take? (Choose 3 correct options)

Select 3 answers
A.Delete the Cloud KMS key after data is loaded to BigQuery
B.Enable CMEK on the Cloud Storage bucket by specifying the KMS key
C.Configure the BigQuery dataset to use a CMEK key
D.Use Google-managed encryption keys
E.Create a key ring and key in Cloud Key Management Service
AnswersB, C, E

Specifying a Cloud KMS key on the Cloud Storage bucket enforces customer-managed encryption at rest for the raw IoT data, satisfying the stem's CMEK constraint at the storage layer. Without this, Google-managed keys apply by default, so the raw landing zone would fail the customer-managed key requirement before BigQuery processing begins.

Why this answer

The scenario requires customer-managed encryption keys (CMEK) for data at rest across Cloud Storage and BigQuery, so the foundational step is option E: create a key ring and key in Cloud Key Management Service, since CMEK always begins with a KMS key ring and a key (or key version) that will be referenced by the services. Option B is correct because enabling CMEK on the Cloud Storage bucket by specifying the KMS key ensures the raw IoT data written to the bucket is encrypted at rest with that customer-managed key rather than a Google-managed key. Option C is correct because configuring the BigQuery dataset to use a CMEK key ensures the processed analytics data stored in BigQuery is also encrypted at rest with the customer-managed key, satisfying the end-to-end requirement.

Option A is wrong because deleting the Cloud KMS key after loading would make the encrypted data unrecoverable and break decryption for both Cloud Storage and BigQuery. Option D is wrong because Google-managed encryption keys are the default and do not meet the requirement for customer-managed keys.

Exam trap

A common trap is thinking that only one component needs CMEK enabled, but for this data pipeline, both Cloud Storage and BigQuery must be configured to use the same customer-managed key. Another trap is selecting 'Delete the Cloud KMS key after data is loaded' (option A), which is incorrect because it would make data unrecoverable and non-compliant.

638
MCQeasy

A startup wants to analyze user clickstream data stored in Cloud Storage in Parquet format. They need to run ad-hoc SQL queries without managing any servers and want to pay only for the queries they run. Which Google Cloud service should they use?

A.BigQuery
B.Dataproc
C.Cloud Bigtable
D.Cloud SQL
AnswerA

BigQuery is a serverless, highly scalable data warehouse that supports querying external data in Cloud Storage via external tables. It charges based on the amount of data scanned per query, aligning with pay-per-query. It requires no infrastructure management, making it ideal for ad-hoc SQL analysis.

Why this answer

BigQuery provides serverless SQL analytics with pay-per-query pricing and can query external Parquet data in Cloud Storage. It requires no infrastructure management, perfectly matching the startup's needs. Other services either require provisioning, are not designed for SQL analytics, or lack external data querying.

Exam trap

The trap here is assuming that any data processing service can query Parquet files, but only BigQuery offers serverless SQL with external table support.

639
MCQmedium

Your team stores IoT sensor readings in a BigQuery table `project.sensors.readings` with columns `sensor_id` (STRING), `reading_time` (TIMESTAMP), and `temperature` (FLOAT64). You need to create a new table that adds a column `avg_temp_7d` containing, for each row, the average temperature of that sensor over the preceding 7 days (including the current row). Which SQL feature should you use to compute this efficiently?

A.A GROUP BY on sensor_id and a date truncation to week, then joining back to the original table
B.A user-defined function (UDF) that loops over the previous 7 days and accumulates temperatures
C.A correlated subquery that selects the average temperature where reading_time is between the current row's time minus 7 days and the current row's time
D.A window function with a RANGE-based frame between 7 days preceding and CURRENT ROW
AnswerD

RANGE BETWEEN INTERVAL 7 DAY PRECEDING AND CURRENT ROW defines a dynamic window based on the timestamp value, so each row's average includes all readings from the same sensor within the preceding 7 days. This is the correct approach for time-based rolling aggregates in BigQuery and avoids self-joins or manual date arithmetic.

Why this answer

A window function with a RANGE frame based on the timestamp column correctly computes a rolling 7-day average per sensor. It is efficient because BigQuery processes the window in a single pass, and the RANGE clause dynamically adjusts the window based on the actual time values, not row counts. This meets the requirement of including the current row and the preceding 7 days.

Exam trap

The trap here is assuming that a ROWS-based window frame (e.g., 7 PRECEDING) works for time-based rolling averages; it would instead include the previous 7 rows regardless of time gaps.

640
MCQeasy

The push endpoint is returning 500 errors. What is the most likely cause?

A.The push endpoint requires authentication but none is set
B.The topic has no messages
C.The push endpoint is not a valid HTTPS URL
D.The ack deadline is too short
AnswerA

If the endpoint expects an Authorization header, requests without it will fail with 500 or 401.

Why this answer

The push endpoint likely requires authentication, but none is configured, causing the 500 errors.

641
MCQmedium

You are designing a Cloud Composer workflow that loads data from Cloud Storage into BigQuery, runs a Dataflow job to transform the data, and then triggers a Dataproc Spark job. After each step, you need to conditionally branch based on success or failure. Which Airflow feature allows you to pass messages between tasks to enable dynamic branching?

A.Sensors
B.XComs
C.TaskFlow API
D.DAG dependencies
AnswerB

XComs let a task push a small value, such as a success flag or branch key, that downstream tasks pull, enabling conditional branching via BranchPythonOperator. This passes messages between tasks without external storage, satisfying the requirement for dynamic branching after each step.

Why this answer

XComs (cross-communications) are Airflow's built-in mechanism for passing small pieces of data between tasks. A task can push a value via xcom_push() or by returning it, and downstream tasks pull it via xcom_pull(), enabling dynamic decisions such as choosing a branch based on a prior task's output. This is exactly what conditional branching in a Composer DAG requires.

Exam trap

PDE often tests the confusion between XComs (data passing) and the TaskFlow API (a coding style that uses XComs) — candidates pick TaskFlow API thinking it is the transport mechanism.

How to eliminate wrong answers

Option A is wrong because Sensors are operators that wait for a condition (a file, a partition, a time) to become true — they do not carry payloads between tasks. Option C is wrong because the TaskFlow API is a decorator-based authoring style (@task) that simplifies DAG code and implicitly uses XComs under the hood; it is not itself the message-passing mechanism. Option D is wrong because DAG dependencies (>>, set_upstream/set_downstream) only define execution order, not data transfer.

642
MCQeasy

Your company uses Cloud Dataflow to process streaming data from Pub/Sub. The pipeline occasionally fails with a 'worker terminated unexpectedly' error. What is the most likely cause of this error?

A.Insufficient memory per worker causing OOM errors
B.Incorrect VPC firewall rules blocking internal communication
C.Staging location bucket lacks write permissions
D.Pub/Sub subscription throughput quota exceeded
AnswerA

Worker termination typically results from the worker process being killed, and out-of-memory conditions are the common trigger: the JVM heap or container exceeds available memory, so the worker dies. Insufficient memory per worker directly explains the unexpected termination.

Why this answer

The 'worker terminated unexpectedly' error in Cloud Dataflow typically indicates that a worker process ran out of memory (OOM) and was killed by the operating system. This occurs when the pipeline's memory requirements exceed the configured worker machine type's memory capacity, often due to large windowing accumulations, skewed data, or inefficient state handling.

Exam trap

Google Cloud often tests the distinction between infrastructure-level errors (like OOM) and configuration or permission errors, so candidates may incorrectly attribute the generic 'worker terminated' message to network or IAM issues rather than resource exhaustion.

How to eliminate wrong answers

Option B is wrong because VPC firewall rules blocking internal communication would cause connectivity errors like 'unable to connect to shuffle service' or 'worker cannot reach Dataflow service', not a generic termination error. Option C is wrong because staging location bucket lacking write permissions would cause a pipeline submission failure with a permission denied error, not a runtime worker termination. Option D is wrong because Pub/Sub subscription throughput quota exceeded would result in Pub/Sub-specific errors such as 'RESOURCE_EXHAUSTED' or backlog buildup, not a worker termination.

643
MCQeasy

A company trains a custom model using TensorFlow and wants to deploy it to Vertex AI for low-latency predictions. The model is large (2 GB). Which deployment option should they choose?

A.Use Vertex AI Batch Prediction job
B.Deploy as a Cloud Function
C.Deploy to Vertex AI Endpoint with a custom container
D.Deploy to Cloud Run with minimum instances
AnswerC

A 2 GB custom TensorFlow model exceeds standard pre-built container limits, so a custom container on a Vertex AI Endpoint is required. This satisfies the low-latency prediction constraint by giving full control over serving runtime and dependencies.

Why this answer

Deploying a large (2 GB) model to Vertex AI Endpoint with a custom container allows you to package the model, its dependencies, and a serving framework (e.g., TensorFlow Serving) into a Docker image. This approach supports low-latency predictions by keeping the model loaded in memory across requests, and it can scale to handle real-time inference traffic, unlike batch or serverless options that have cold-start or size limitations.

Exam trap

Google Cloud often tests the misconception that Cloud Run or Cloud Functions can handle large models for real-time inference, ignoring their size limits, cold-start latency, and lack of native Vertex AI integration for model management and scaling.

How to eliminate wrong answers

Option A is wrong because Vertex AI Batch Prediction is designed for asynchronous, high-throughput processing of large datasets, not for low-latency real-time predictions; it processes jobs in batches and does not maintain a persistent endpoint. Option B is wrong because Cloud Functions have a maximum deployment size of 2 GB (unpackaged) and a 60-second timeout, making them unsuitable for a 2 GB model that requires persistent memory and low-latency inference. Option D is wrong because Cloud Run has a container image size limit of 2 GB (uncompressed) and a request timeout of 60 minutes, but it lacks native integration with Vertex AI's model registry and optimized serving infrastructure, and it may incur cold-start latency even with minimum instances.

644
MCQhard

A healthcare company stores patient records in a Cloud Storage bucket that must remain in a specific region for data residency. The security team requires that all data be encrypted with keys the company controls and can rotate, and that access to the keys be auditable. The company also wants to avoid managing key material on-premises. Which approach should the data engineer choose?

A.Use Customer-Managed Encryption Keys in Cloud KMS with a key ring in the required region.
B.Use Customer-Supplied Encryption Keys and store the key material in a local secrets file.
C.Use Google-managed encryption keys and rely on Google's default encryption for the bucket.
D.Use Cloud External Key Manager to connect to a third-party external key management system.
AnswerA

CMEK in Cloud KMS lets the company own, rotate, and disable keys while Google manages the underlying infrastructure, satisfying the no-on-premises requirement. A regional key ring keeps key material aligned with the data residency constraint, and Cloud KMS logs key operations to Cloud Audit Logs for auditability. This meets all stated requirements.

Why this answer

Customer-Managed Encryption Keys in Cloud KMS give the company ownership and rotation control while Google operates the key infrastructure, and a regional key ring aligns key storage with the data residency requirement. Cloud KMS also writes key usage to audit logs. Default encryption, customer-supplied keys, and external key managers either remove control or push key management outside Google Cloud.

Exam trap

The trap here is equating default Google-managed encryption with customer-controlled keys, when only CMEK provides rotation and auditable key ownership.

645
MCQeasy

A developer wants to create a BigQuery table that automatically expires data older than 30 days to reduce storage costs. Which table design feature should be used?

A.Authorized view
B.Clustered table
C.Materialized view
D.Partitioned table with partition expiration
AnswerD

Partition expiration automatically deletes partitions once their partition date passes the specified retention period, so data older than 30 days is removed without manual jobs. This satisfies the automatic expiry requirement while keeping recent data queryable, unlike table-level expiration which drops the entire table.

Why this answer

BigQuery partitioned tables support partition expiration, which automatically deletes partitions older than a specified number of days. Setting a partition expiration of 30 days on a time-partitioned table will drop data older than 30 days, reducing storage costs. This is the native, cost-effective way to enforce data retention.

Exam trap

PDE often tests the difference between partitioning, clustering, and expiration features, and candidates may confuse clustering with retention or pick materialized views for cost reduction.

How to eliminate wrong answers

Option A is wrong because an authorized view controls access to data, not retention or expiration. Option B is wrong because clustering only sorts data within partitions to improve query performance; it does not expire data. Option C is wrong because a materialized view precomputes query results for performance, not for data lifecycle management.

646
MCQmedium

A company uses BigQuery for analytics. They have a table that is queried frequently by date range. To reduce costs, they want to ensure queries only scan the relevant partitions. They also want to improve performance for queries filtering on a specific customer_id. Which table design should they use?

A.Partition by ingestion time and cluster by customer_id
B.Use a materialized view that filters by date and customer_id
C.Cluster by date column and partition by customer_id
D.Partition by date column and cluster by customer_id
AnswerD

Partitioning on the date column restricts each query to the relevant date range, cutting bytes scanned and cost. Clustering by customer_id then sorts data within each partition, so filters on that column skip irrelevant blocks, satisfying both the cost and performance constraints.

Why this answer

To reduce costs by ensuring queries only scan relevant partitions, you should partition by the date column (so date-range filters prune partitions). To improve performance for queries filtering on customer_id, you should cluster by customer_id (so BigQuery can skip blocks within partitions). Therefore, partition by date column and cluster by customer_id is the correct design.

Exam trap

PDE often tests the distinction between partitioning and clustering — candidates may reverse them or choose ingestion-time partitioning when a business date column is the filter, leading to unnecessary data scans.

How to eliminate wrong answers

Option A is wrong because partitioning by ingestion time does not align with date-range queries on a business date column; ingestion time may not match the query filter, leading to full scans. Option B is wrong because a materialized view can improve performance but does not change the underlying table design for cost reduction; it also adds storage cost and does not address clustering for customer_id. Option C is wrong because clustering by date and partitioning by customer_id reverses the optimal design — partitioning by customer_id would create too many partitions and date-range queries would not benefit from partition pruning.

647
MCQhard

A retail company stores 500 TB of JSON transaction logs in Cloud Storage. Analysts need to run ad hoc SQL over the data, and the company wants to avoid managing a separate cluster while keeping query cost predictable. The logs are already partitioned into date-based prefixes. What should the data engineer do?

A.Create a Dataproc cluster with Hive and query the JSON logs through Hive external tables on Cloud Storage.
B.Load all 500 TB into a BigQuery native table partitioned by ingestion time and query the native table.
C.Create a BigQuery external table over the Cloud Storage prefixes and use a hive-partitioned layout so queries use partition pruning.
D.Use Cloud SQL for PostgreSQL with the foreign data wrapper to read the JSON files from Cloud Storage.
AnswerC

BigQuery external tables can query Cloud Storage data directly without loading it and without a cluster to manage. Defining hive partitioning on the date prefixes lets BigQuery prune partitions based on the date filter, which keeps query cost tied to the partitions scanned rather than the whole 500 TB.

Why this answer

An external table over Cloud Storage gives serverless SQL access to the JSON logs, and hive partitioning on the date prefixes enables pruning so only the relevant date partitions are read. This avoids loading 500 TB, removes cluster management, and keeps cost predictable because bytes billed track the partitions scanned.

Exam trap

The trap here is assuming external tables always scan all files, when hive partitioning lets BigQuery prune based on the partition key in the path.

648
MCQmedium

A retail company ingests point-of-sale events from thousands of stores into Cloud Pub/Sub. They need to process these events in a streaming Dataflow pipeline that enriches each event with store metadata from a slowly changing BigQuery table. The enrichment table is updated only once per day. The pipeline must minimize latency and avoid querying BigQuery for every event. Which approach should they use?

A.Use a CoGroupByKey transform to join the streaming events with a bounded read of the metadata table.
B.Load the BigQuery metadata table into a side input as a key-value map, refreshing it periodically via a scheduled pipeline.
C.Configure the pipeline to call the BigQuery Storage Read API for each event to fetch the latest metadata.
D.Use BigQueryIO.Read with a query that joins the events to the metadata table, and apply the join in the pipeline.
AnswerB

Side inputs in Dataflow allow you to supply additional data to each element in a pipeline without querying an external system per element. By loading the slowly changing metadata into a side input and refreshing it periodically, the pipeline can enrich events with minimal latency and without hitting BigQuery for each event. This matches the requirement to minimize latency and avoid per-event queries.

Why this answer

The correct approach is to load the slowly changing metadata into a side input that is refreshed periodically. This allows the streaming pipeline to enrich each event without querying BigQuery per event, keeping latency low. Side inputs are a standard Dataflow pattern for enriching streaming data with reference data that changes infrequently, and they avoid the high cost and latency of per-element external calls.

Exam trap

The trap here is assuming that a streaming pipeline must query BigQuery for each event or use a join transform, when side inputs are designed for exactly this kind of enrichment.

649
MCQhard

A company processes financial transactions using Cloud Dataflow. They need to ensure that late-arriving data is handled correctly for fraud detection. The pipeline uses event time processing. Which approach should they use to handle late data?

A.Sliding windows with early firing
B.Session windows with gap duration
C.Fixed windows with allowed lateness
D.Global windows with triggers
AnswerC

Allowed lateness extends a fixed window's lifetime, so late-arriving events still trigger updated panes rather than being discarded as dropped data. This satisfies the event-time fraud-detection requirement by emitting revised results after the watermark passes.

Why this answer

Fixed windows with allowed lateness are the standard approach in Cloud Dataflow (Apache Beam) for handling late-arriving data in event-time processing. By specifying an allowed lateness duration, the pipeline retains the window state for that period, allowing late events to be correctly assigned to their original window and triggering recomputation of results. This ensures fraud detection pipelines can account for delayed transactions without missing or misordering data.

Exam trap

Google Cloud often tests the misconception that sliding or session windows inherently handle late data, when in fact only explicit allowed lateness (or a similar mechanism) provides the necessary state retention and watermark adjustment for late-arriving events.

How to eliminate wrong answers

Option A is wrong because sliding windows with early firing are designed to produce speculative results before the window closes, not to handle late-arriving data; early firing does not extend the window to accept late events. Option B is wrong because session windows with gap duration are used to group events into sessions based on inactivity gaps, not to manage late data; they do not provide a mechanism to accept events that arrive after the session has closed. Option D is wrong because global windows with triggers are typically used for unbounded aggregations where all data belongs to a single window, but they do not naturally handle late-arriving data within specific time boundaries required for fraud detection; they lack the per-window lateness cutoff that fixed windows offer.

650
Multi-Selectmedium

A company needs to stream real-time user activity data from their application into BigQuery for immediate dashboarding. They want to minimize latency (under 5 seconds) and ensure exactly-once delivery. Which TWO options should they consider? (Choose 2)

Select 2 answers
A.Use Cloud Functions to receive events and call the BigQuery REST API
B.Use BigQuery Storage Write API in committed mode
C.Use BigQuery legacy streaming inserts directly from the application
D.Use Apache Kafka on Dataproc and write to BigQuery via the BigQuery Kafka connector
E.Stream data to Pub/Sub, then use Dataflow to write to BigQuery with exactly-once guarantees
AnswersB, E

BigQuery Storage Write API in committed mode provides exactly-once semantics through stream-level offsets, so duplicate records are discarded on retry. It supports sub-second, real-time ingestion directly into BigQuery, satisfying the under-five-second latency requirement without intermediate staging. This makes it ideal for streaming live user activity into dashboards.

Why this answer

Option B is correct because the BigQuery Storage Write API in committed mode provides exactly-once semantics via write streams and offsets, and it supports low-latency streaming ingestion suitable for sub-5-second dashboarding. Option E is correct because Pub/Sub plus Dataflow is the canonical Google Cloud streaming pipeline: Pub/Sub ingests events durably, and Dataflow's BigQueryIO in STREAMING mode with exactly-once processing guarantees deduplication and exactly-once writes to BigQuery. Option A is not ideal because Cloud Functions invoking the BigQuery REST API (tabledata.insertAll) does not provide exactly-once delivery and adds per-event overhead and cold-start latency.

Option C is wrong because legacy streaming inserts offer at-least-once semantics and can produce duplicate rows, violating the exactly-once requirement. Option D is not the best fit because running Kafka on Dataproc adds operational complexity, and the Kafka connector typically relies on the Storage Write API or streaming inserts without inherently guaranteeing exactly-once end-to-end delivery in this scenario.

Exam trap

Google Cloud often tests the distinction between legacy streaming inserts (at-least-once) and the Storage Write API in committed mode (exactly-once). Candidates may mistakenly choose legacy inserts because they are simpler to implement, ignoring the exactly-once requirement.

651
MCQeasy

A data engineer needs to automatically delete objects from a Cloud Storage bucket after 30 days and archive them to nearline storage after 7 days. Which configuration should they use?

A.Set a lifecycle rule to SetStorageClass to nearline after 30 days only
B.Set a lifecycle rule to delete objects after 7 days only
C.Set a lifecycle rule to SetStorageClass to nearline after 7 days and delete after 30 days
D.Set a lifecycle rule to delete objects after 7 days and SetStorageClass to nearline after 30 days
AnswerC

A single lifecycle rule can hold multiple conditional actions. SetStorageClass transitions objects to Nearline after 7 days, then the delete action removes them at 30 days, matching both the archival and deletion constraints in one configuration.

Why this answer

It implements a lifecycle rule that first transitions objects to Nearline storage after 7 days (reducing costs for infrequently accessed data) and then deletes them after 30 days. This matches the requirement to archive after 7 days and delete after 30 days, using the `SetStorageClass` and `Delete` actions in the correct chronological order.

Exam trap

Google Cloud often tests the order of lifecycle actions: candidates mistakenly think deletion should come before archiving, but the correct sequence is to archive first (to reduce cost) and delete later, as objects cannot be archived after deletion.

How to eliminate wrong answers

Option A is wrong because it only sets the storage class to Nearline after 30 days, missing the deletion requirement entirely and incorrectly archiving after 30 days instead of 7. Option B is wrong because it only deletes objects after 7 days, ignoring the archive-to-Nearline step and deleting data too early. Option D is wrong because it reverses the order: it deletes objects after 7 days (before they can be archived) and then attempts to set storage class to Nearline after 30 days, which is impossible since the objects are already deleted.

652
MCQmedium

A company wants to automate model retraining and deployment whenever new training data becomes available. Which service should be used to orchestrate the end-to-end workflow?

A.Cloud Build
B.Vertex AI Pipelines
C.Cloud Scheduler
D.Cloud Composer
AnswerB

Vertex AI Pipelines orchestrates the full retraining-to-deployment workflow, triggered when new training data arrives. It chains data preprocessing, training, evaluation and deployment as managed pipeline steps, satisfying the stem's requirement to automate the end-to-end process rather than run isolated jobs.

Why this answer

Vertex AI Pipelines is the correct choice because it is a managed service specifically designed to orchestrate and automate end-to-end ML workflows, including model retraining and deployment triggered by new data. It allows you to define pipelines as a directed acyclic graph (DAG) of steps using the Kubeflow Pipelines SDK or pre-built components, and it integrates natively with other Vertex AI services for training, evaluation, and deployment.

Exam trap

The trap here is that candidates often confuse Cloud Composer (a general-purpose Airflow service) with Vertex AI Pipelines, but the exam expects you to recognize that Vertex AI Pipelines is the ML-specific, fully managed solution for end-to-end ML workflow orchestration, while Cloud Composer requires more manual setup and lacks native Vertex AI integration.

How to eliminate wrong answers

Option A is wrong because Cloud Build is a CI/CD service focused on building, testing, and deploying software artifacts (e.g., container images), not on orchestrating ML workflows with steps like data validation, model training, and deployment. Option C is wrong because Cloud Scheduler is a cron job service that triggers actions on a time-based schedule, not on the event of new training data becoming available, and it lacks the workflow orchestration capabilities needed for complex ML pipelines. Option D is wrong because Cloud Composer is a managed Apache Airflow service that can orchestrate workflows, but it is a general-purpose workflow orchestrator, not purpose-built for ML pipelines; Vertex AI Pipelines provides tighter integration with Vertex AI components, managed execution, and artifact tracking, making it the more appropriate choice for this specific ML automation scenario.

653
MCQhard

You are optimizing a BigQuery query that runs on a large table (hundreds of TB). The table is partitioned by date and frequently queried with filters on a specific customer_id column and date range. Queries are slow even after partitioning. Which optimization should you apply?

A.Increase the number of BigQuery slots
B.Columnar clustering on customer_id
C.Create materialized views for each customer
D.Denormalize the table to reduce joins
AnswerB

Clustering sorts and co-locates data by customer_id within each date partition, so filtered scans read far fewer blocks. Partitioning alone cannot prune on customer_id, which is why queries stay slow; clustering on that column satisfies the filter constraint.

Why this answer

Clustering on customer_id within the partition improves query performance because BigQuery can prune blocks based on clustered columns. Partitioning alone doesn't help with non-date filters. Materialized views may help pre-aggregated queries but not ad-hoc customer_id filters.

Denormalization is not an optimization. Increasing slots is expensive and doesn't address data structure.

654
MCQeasy

A company uses Dataflow to process streaming data from Pub/Sub. They notice increased processing latency. What is the most likely cause?

A.Insufficient workers
B.Pub/Sub subscription issue
C.Too many shards
D.Wrong machine type
AnswerA

Insufficient workers create backpressure and increased latency as the pipeline cannot keep up with throughput.

Why this answer

In Dataflow, processing latency increases most commonly due to insufficient workers, as the streaming pipeline cannot keep up with the incoming data rate when the number of Compute Engine instances is too low. This causes backpressure from Pub/Sub, leading to growing unacknowledged messages and higher end-to-end latency. Autoscaling may be delayed or limited by max worker count settings, making manual or configuration-based worker scaling the primary corrective action.

Exam trap

Google Cloud often tests the misconception that Pub/Sub subscription issues (like ack deadline) are the primary cause of latency, but the trap here is that latency in Dataflow is almost always a worker scaling problem, not a Pub/Sub configuration issue.

How to eliminate wrong answers

Option B is wrong because a Pub/Sub subscription issue (e.g., expired pull request or misconfigured ack deadline) would cause message delivery failures or duplicates, not a gradual increase in processing latency across the pipeline. Option C is wrong because too many shards (i.e., excessive parallelism) can cause overhead but typically leads to underutilization or increased cost, not increased latency; latency from too many shards is rare and usually secondary to worker count. Option D is wrong because the wrong machine type (e.g., low CPU or memory) could degrade per-worker performance, but the most likely and direct cause of increased latency in a streaming Dataflow job is insufficient worker count, not machine type, as Dataflow’s autoscaling primarily adjusts worker count rather than machine type.

655
MCQeasy

A retail company wants to analyze point-of-sale transaction data stored in Cloud SQL for PostgreSQL. They need to run complex analytical queries joining this data with data in BigQuery. The data changes frequently, and they want near-real-time access without impacting the production Cloud SQL instance. Which approach should they use?

A.Export Cloud SQL data to Cloud Storage daily and load into BigQuery
B.Use BigQuery federated queries to query Cloud SQL directly
C.Use Cloud Data Fusion to replicate Cloud SQL to BigQuery with a batch pipeline
D.Set up a Datastream stream from Cloud SQL to BigQuery
AnswerD

Datastream provides change data capture (CDC) from Cloud SQL for PostgreSQL to BigQuery, replicating changes in near-real-time. It reads from the source's replication log without impacting production query performance and writes to BigQuery with low latency. This enables complex analytical joins in BigQuery while keeping the production database isolated.

Why this answer

Datastream is a serverless change data capture service that replicates data from Cloud SQL for PostgreSQL to BigQuery in near-real-time without impacting the source database's performance. It reads the write-ahead log and streams changes, enabling up-to-date analytics. Other options either introduce latency, impact production, or require manual intervention.

Exam trap

The trap here is choosing federated queries because they seem to provide direct access, but they can impact production and do not synchronize data continuously.

656
MCQhard

A data engineer needs to split time-series data for training a forecasting model. The data is sorted by timestamp. The engineer wants to avoid leakage where future data influences training. Which data splitting approach should they use?

A.Use k-fold cross-validation with random assignment
B.Use stratified splitting on the target variable
C.Perform a random 80/20 split on the entire dataset
D.Use a time-series aware split: first 80% of data by timestamp for training, last 20% for testing
AnswerD

Splitting chronologically by timestamp keeps all training observations earlier than every test observation, so the model never learns from future values. Random splitting would leak future information backwards into training, violating the no-leakage constraint for forecasting time-series data.

Why this answer

Time-series data has a temporal order, so training must only use data that precedes the test data to prevent future information from leaking into the model. Option D holds out the last 20% of records by timestamp for testing and trains on the earlier 80%, which mirrors real forecasting conditions where the model predicts unseen future values. This preserves causality and gives a realistic estimate of out-of-sample performance.

Exam trap

The trap here is that candidates reflexively choose k-fold cross-validation because it is the default best practice for i.i.d. data, forgetting that temporal ordering invalidates random shuffling.

How to eliminate wrong answers

Option A is wrong because k-fold cross-validation with random assignment shuffles observations across folds, so the model can be trained on future timestamps and tested on past ones, directly causing temporal leakage. Option B is wrong because stratified splitting on the target variable preserves class proportions but ignores timestamp ordering, so future data can still end up in the training set. Option C is wrong because a random 80/20 split also ignores the temporal order and allows future observations into training, producing optimistically biased metrics.

657
Multi-Selectmedium

A data team is building a near-real-time dashboard that displays aggregated metrics from Kafka topics. They want to use Pub/Sub as a managed messaging service and Dataflow for stream processing. They need to ingest data from Kafka into Pub/Sub with minimal custom code. Which THREE Google Cloud services should they use together? (Choose three.)

Select 3 answers
A.Dataflow
B.Pub/Sub
C.Kafka Connect (with Pub/Sub connector)
D.Cloud NAT
E.Cloud Functions
AnswersA, B, C

Dataflow provides the managed Apache Beam runtime that reads from Kafka and writes into Pub/Sub, satisfying the minimal-custom-code constraint. Its built-in Kafka-to-Pub/Sub template performs the ingestion without bespoke connectors, then continues stream processing for the dashboard's aggregated metrics.

Why this answer

Option A, Dataflow, is correct because it is Google Cloud's managed Apache Beam service for stream processing, and it can run a streaming pipeline that reads from Pub/Sub, performs windowed aggregations, and writes results for the near-real-time dashboard. Option B, Pub/Sub, is correct because it serves as the managed messaging service that decouples the Kafka ingestion layer from the Dataflow processing layer, buffering messages and enabling reliable, scalable delivery. Option C, Kafka Connect (with Pub/Sub connector), is correct because Kafka Connect provides a configuration-driven, low-code way to move data from Kafka topics into Pub/Sub using a Pub/Sub sink connector, satisfying the requirement for minimal custom code.

Option D, Cloud NAT, is not relevant because it provides outbound internet address translation for private VMs and does not ingest Kafka data into Pub/Sub. Option E, Cloud Functions, is not appropriate here because it is an event-driven serverless compute service for lightweight functions, not the managed stream-processing engine needed for aggregated Kafka metrics.

658
MCQhard

A company runs a critical real-time data pipeline using Dataflow that ingests events from Cloud Pub/Sub, performs aggregations using sliding windows, and writes results to BigQuery. The pipeline is deployed in us-central1. The pipeline's latency has increased recently, and the Dataflow monitoring shows that the 'system lag' metric is consistently above 5 minutes. The pipeline is using Streaming Engine and has 10 workers with 4 vCPUs each. The pipeline processes approximately 100,000 events per second. The team has verified that the source Pub/Sub topic has sufficient publish throughput and the BigQuery table has no quota issues. The pipeline logs show that some workers are experiencing GC overhead limit exceeded errors. The pipeline code uses stateful processing with a custom keyed state for deduplication. What is the most likely cause of the increased latency?

A.The number of workers is insufficient; increasing to 20 workers will reduce latency.
B.The stateful processing is causing large state sizes that lead to GC overhead; use a more efficient state backend or increase worker memory.
C.The sliding window duration is too long; reducing it to 1 minute will improve performance.
D.The deduplication logic is causing a bottleneck; removing it will reduce latency.
AnswerB

Keyed deduplication state grows unbounded per key, inflating worker heap and triggering GC overhead limit errors that stall processing and raise system lag. Switching to a more efficient state backend or adding worker memory directly relieves the GC pressure causing the latency.

Why this answer

The GC overhead limit exceeded errors indicate that workers are spending too much time garbage collecting, which is a classic symptom of excessive heap memory usage. Stateful processing with custom keyed state for deduplication can cause large per-key state sizes, especially with sliding windows that maintain overlapping state for each key. This forces the JVM to constantly garbage collect, increasing system lag beyond 5 minutes.

Using a more efficient state backend (e.g., reducing state size or using Dataflow's built-in deduplication) or increasing worker memory directly addresses the root cause.

Exam trap

Google Cloud often tests the misconception that scaling workers (Option A) is the universal fix for latency, when in reality memory-related issues like GC overhead require tuning state management or worker resources, not just parallelism.

How to eliminate wrong answers

Option A is wrong because increasing the number of workers does not fix the GC overhead issue; it may even worsen it by distributing state across more workers without reducing per-worker memory pressure. Option C is wrong because reducing the sliding window duration does not address the state size or GC problem; it could actually increase the number of overlapping windows and state churn. Option D is wrong because removing deduplication would compromise data correctness; the bottleneck is not the logic itself but the memory footprint of the state, which can be mitigated without removing the feature.

659
MCQmedium

A company is deploying a large-scale streaming application on Google Kubernetes Engine. They need to ensure the application can handle sudden traffic spikes without dropping data. Which architectural pattern is most appropriate?

A.Implement custom retry logic with exponential backoff in the application.
B.Use Cloud SQL as a temporary buffer and process from there.
C.Pre-provision 3x the expected peak capacity to handle spikes.
D.Use a Pub/Sub topic as a buffer and autoscale consumer pods based on Pub/Sub subscription backlog.
AnswerD

Pub/Sub decouples ingestion from processing, absorbing traffic spikes as a durable buffer so no data is dropped. Autoscaling consumer pods on subscription backlog matches capacity to demand, directly satisfying the stem's requirement to handle sudden spikes without loss.

Why this answer

Pub/Sub provides a durable, scalable, and asynchronous message buffer that decouples the producer from the consumer. By autoscaling consumer pods based on the Pub/Sub subscription backlog (e.g., using the 'pubsub.googleapis.com/subscription/num_undelivered_messages' custom metric with Horizontal Pod Autoscaler), the application can elastically handle traffic spikes without data loss, as messages are persisted until acknowledged.

Exam trap

The trap here is that candidates confuse buffering with retry logic or database storage, failing to recognize that Pub/Sub is the Google Cloud-native service specifically designed for decoupling and buffering in event-driven architectures.

How to eliminate wrong answers

Option A is wrong because custom retry logic with exponential backoff addresses transient failures but does not provide a buffer for sudden traffic spikes; if the producer outpaces the consumer, data is still dropped or rejected. Option B is wrong because Cloud SQL is not designed as a message buffer; it is a relational database with limited throughput and connection scaling, and using it as a temporary buffer would create a bottleneck and risk data loss under high load. Option C is wrong because pre-provisioning 3x the expected peak capacity leads to significant cost overprovisioning and still cannot guarantee handling of unexpected spikes beyond that factor; it violates the cloud-native principle of elastic scaling.

660
MCQmedium

A logistics company uploads millions of small JSON files per day to a Cloud Storage bucket and needs to query them with standard SQL immediately, but the analytics team does not want to create BigQuery tables or manage schema changes. The files follow a consistent structure but new fields are added frequently. Which BigQuery capability should they use to query these objects directly from Cloud Storage with the least operational overhead?

A.Use the BigQuery Storage Write API to stream the JSON files into a table.
B.Create an external table over the Cloud Storage bucket and query it with standard SQL.
C.Mount the Cloud Storage bucket as a BigQuery dataset and query the files as tables.
D.Load the JSON files into a BigQuery table using a recurring Dataflow batch job.
AnswerB

BigQuery external tables let you query data stored in Cloud Storage with standard SQL without loading it, and you can define the schema or let BigQuery infer it. This matches the requirement to avoid managing loads and table schema changes, though query performance is lower than native tables and you are billed for bytes scanned in Cloud Storage.

Why this answer

An external table defined over Cloud Storage allows standard SQL queries against the JSON objects without loading them, and schema autodetection handles evolving fields. This minimizes operational overhead because there is no pipeline or table to maintain, at the cost of query performance and per-byte scanning charges on the external data.

Exam trap

The trap here is assuming that querying Cloud Storage data requires loading it into BigQuery first; external tables exist precisely to query objects in place.

661
MCQhard

A company's Dataflow pipeline uses the PubSubIO source to read messages and writes to BigQuery via the BigQueryIO sink. The pipeline is running in Streaming mode with exactly-once semantics enabled. Occasionally, duplicate rows appear in BigQuery. What is the most likely reason?

A.The user-provided record ID for deduplication in BigQuery's streaming inserts is not being set for all messages, leading to duplicate rows.
B.The pipeline is using the WriteResult method with WRITE_APPEND in batch mode, which can cause duplicates if retries happen.
C.The pipeline is experiencing the 'dataflow streaming log processing' bug, causing duplicate logs to be written.
D.The PubSubIO source is configured with a dead-letter queue and messages are being redelivered without proper deduplication.
AnswerA

BigQueryIO's streaming exactly-once deduplication relies on a deterministic record ID per message; when that ID is absent or non-unique, BigQuery cannot deduplicate retried inserts, so duplicates appear. The pipeline's exactly-once guarantee depends on this ID being set for every message.

Why this answer

In Dataflow streaming pipelines with exactly-once semantics, BigQuery's streaming inserts use user-provided record IDs for deduplication. If the record ID is not set for all messages, BigQuery cannot identify duplicates, and retries or redeliveries from Pub/Sub can result in duplicate rows. This is the most common cause of duplicates in this scenario.

Exam trap

Google Cloud often tests the misconception that exactly-once semantics in Dataflow automatically deduplicates at the sink, but in reality, BigQuery requires explicit user-provided record IDs for deduplication during streaming inserts.

How to eliminate wrong answers

Option B is wrong because WRITE_APPEND in batch mode is not relevant to a streaming pipeline with exactly-once semantics; the question specifies streaming mode, and batch mode duplicates would not explain streaming-specific behavior. Option C is wrong because there is no known 'dataflow streaming log processing' bug that causes duplicate logs; this is a fabricated term. Option D is wrong because a dead-letter queue handles failed messages after retries are exhausted, not redelivery; Pub/Sub redelivery without deduplication is already addressed by the user-provided record ID mechanism, and the dead-letter queue does not cause duplicates.

662
MCQmedium

A company has a trained model stored in Vertex AI Model Registry. They want to automate retraining when new training data arrives in Cloud Storage. Which approach is most efficient?

A.Use Cloud Functions triggered by Cloud Storage events to start a Vertex AI Training job
B.Use Dataflow to continuously update the model
C.Use Cloud Scheduler to trigger a Cloud Build retraining step
D.Schedule a weekly Cloud Composer DAG to check for new data and retrain
AnswerA

A Cloud Function triggered by Cloud Storage object-finalise events reacts immediately to new training data, programmatically launching a Vertex AI Training job. This event-driven pattern satisfies the automation constraint without polling, and is more efficient than scheduled retraining or manual pipeline runs.

Why this answer

Cloud Functions can be directly triggered by Cloud Storage events (e.g., object finalize) to invoke the Vertex AI Training service via the AI Platform API. This creates an event-driven, serverless pipeline that retrains the model immediately when new data arrives, without polling or manual intervention, making it the most efficient and cost-effective approach.

Exam trap

Google Cloud often tests the distinction between event-driven (Cloud Functions) and scheduled (Cloud Scheduler, Cloud Composer) approaches, and candidates mistakenly choose a scheduled option thinking it is simpler, missing the requirement for immediate reaction to new data.

How to eliminate wrong answers

Option B is wrong because Dataflow is a stream/batch data processing service for transforming data, not for orchestrating model retraining; it would require custom code to trigger training and lacks native integration with Vertex AI Model Registry. Option C is wrong because Cloud Scheduler triggers jobs on a fixed schedule, not on data arrival events, so it cannot react to new data in real time and may waste resources on unnecessary retraining. Option D is wrong because a weekly Cloud Composer DAG introduces latency (up to a week) and operational overhead for a simple event-driven task, and it is less efficient than a serverless function that fires instantly on data arrival.

663
MCQhard

You have a BigQuery table 'events' with a TIMESTAMP column 'event_time'. You need to compute, for each event, the difference in seconds from the previous event of the same user. Which window function should you use?

A.FIRST_VALUE(event_time) OVER (PARTITION BY user_id ORDER BY event_time)
B.LEAD(event_time) OVER (PARTITION BY user_id ORDER BY event_time)
C.LAG(event_time) OVER (PARTITION BY user_id ORDER BY event_time)
D.ROW_NUMBER() OVER (PARTITION BY user_id ORDER BY event_time)
AnswerC

LAG(event_time) OVER (PARTITION BY user_id ORDER BY event_time) retrieves the preceding row's timestamp within each user's partition, satisfying the per-user sequential comparison the stem requires. Subtracting it from the current event_time yields the seconds elapsed since that user's previous event, without collapsing rows as aggregation would.

Why this answer

LAG(event_time) OVER (PARTITION BY user_id ORDER BY event_time) returns the event_time value from the previous row within the same user_id partition, ordered chronologically. Subtracting that returned timestamp from the current row's event_time (e.g., TIMESTAMP_DIFF(event_time, LAG(...), SECOND)) yields the seconds elapsed since the user's prior event. This is the canonical pattern for gap/delta calculations in BigQuery analytic functions.

Exam trap

The trap here is confusing LAG with LEAD — candidates often pick LEAD because they think 'previous' maps to the next row, but LAG looks backward and LEAD looks forward.

How to eliminate wrong answers

Option A is wrong because FIRST_VALUE returns the earliest event_time in the partition for every row, not the immediately preceding event, so it cannot produce a per-event delta. Option B is wrong because LEAD looks forward to the next row, which would compute time-to-next-event rather than time-since-previous-event. Option D is wrong because ROW_NUMBER only assigns a sequential integer rank; it carries no timestamp value and therefore cannot be subtracted to produce a time difference.

664
MCQhard

A company uses Vertex AI Pipelines to orchestrate ML workflows. They want to automatically retrain the model when new data arrives, but only if the model's performance drops below a threshold. Which approach is best?

A.Use BigQuery scheduled queries to trigger pipeline
B.Trigger a pipeline on a schedule
C.Use Vertex AI Model Monitor to detect skew and trigger retraining
D.Use Cloud Functions to evaluate performance and trigger pipeline
AnswerC

Model Monitor detects training-serving skew and drift against a baseline, emitting alerts when distributions diverge. Wiring those alerts to a pipeline trigger satisfies the requirement to retrain only when performance degrades, rather than on every new data arrival.

Why this answer

Vertex AI Model Monitor is specifically designed to detect prediction drift and data skew in deployed models. When the monitor identifies that model performance has dropped below a defined threshold, it can automatically trigger a retraining pipeline via a Cloud Function or Pub/Sub notification, ensuring retraining occurs only when necessary rather than on a fixed schedule.

Exam trap

Google often tests the distinction between scheduled retraining (Option B) and event-driven retraining triggered by actual model degradation (Option C), where candidates mistakenly choose a schedule-based approach because they overlook the requirement to retrain 'only if' performance drops below a threshold.

How to eliminate wrong answers

Option A is wrong because BigQuery scheduled queries are used for running SQL queries on a schedule, not for triggering ML pipelines based on model performance metrics. Option B is wrong because triggering a pipeline on a schedule would retrain the model at fixed intervals regardless of whether performance has degraded, wasting resources and potentially deploying unnecessary model versions. Option D is wrong because while Cloud Functions can evaluate performance and trigger a pipeline, this approach requires custom code to monitor model performance and lacks the built-in skew/drift detection capabilities that Vertex AI Model Monitor provides out-of-the-box.

665
MCQeasy

A startup wants to build a data lake on Google Cloud to store raw JSON, CSV, and Parquet files from various sources. They need a storage solution that is highly durable, globally accessible, and integrates natively with BigQuery and Dataproc. They want to minimize management overhead. Which Google Cloud service should they use?

A.Filestore
B.Bigtable
C.Cloud Storage
D.Cloud SQL
AnswerC

Cloud Storage is a highly durable object store that integrates natively with BigQuery external tables and Dataproc. It requires no provisioning and scales automatically, making it ideal for a data lake with minimal management. It supports all mentioned file formats and is globally accessible.

Why this answer

Cloud Storage is the foundational object store for data lakes on Google Cloud. It offers high durability, global accessibility, and native integration with analytics services like BigQuery and Dataproc. It requires no capacity planning or server management, aligning with the goal of minimal overhead.

Relational, NoSQL, and NFS services are not suited for raw file storage at scale.

Exam trap

The trap here is assuming that any storage service can serve as a data lake; only object storage like Cloud Storage provides the necessary scale, durability, and native analytics integration.

666
MCQeasy

A data engineer needs to transfer 500 TB of on-premises data to Google Cloud Storage. The data is stored on NAS devices and the network bandwidth is limited to 100 Mbps. What is the most cost-effective and timely transfer method?

A.Use Storage Transfer Service over the internet
B.Use a VPN connection and rsync
C.Use gsutil cp in parallel
D.Use Transfer Appliance
AnswerD

Transfer Appliance bypasses the 100 Mbps network constraint entirely by shipping physical hardware, making 500 TB feasible in days rather than the months a 100 Mbps link would require. For NAS-resident data at this volume, it is also more cost-effective than sustained egress or Interconnect charges.

Why this answer

Transfer Appliance is Google's physical data transfer service that ships a rackable appliance to the customer site, where data is copied locally and the appliance is shipped back to Google for upload to Cloud Storage. For 500 TB over a 100 Mbps link, the network transfer time would be roughly 1.5 years, making online transfer impractical. Transfer Appliance is the most cost-effective and timely option for this volume and bandwidth.

Exam trap

The trap is underestimating transfer time — candidates pick Storage Transfer Service or gsutil because they are 'cloud-native,' forgetting that 500 TB over 100 Mbps is physically impossible in a reasonable timeframe.

How to eliminate wrong answers

Option A is wrong because Storage Transfer Service over the internet is still bounded by the 100 Mbps link, so 500 TB would take over a year and incur significant egress/transfer costs. Option B is wrong because a VPN plus rsync is even slower due to encryption overhead and is not designed for petabyte-scale bulk migration. Option C is wrong because gsutil cp in parallel only improves throughput marginally and cannot overcome a 100 Mbps physical link limit.

667
MCQhard

A data engineer is designing a batch ETL pipeline that reads CSV files from Cloud Storage, transforms them using Dataproc, and writes the results to BigQuery. The data volume is expected to grow 10x in the next year. Which design approach best balances cost and performance?

A.Create a single large persistent Dataproc cluster to handle the peak load.
B.Use Cloud Data Fusion to visually design the pipeline and run it on Dataproc.
C.Use a Dataproc cluster with preemptible worker nodes and autoscaling enabled.
D.Migrate the pipeline to Dataflow with Apache Beam and use flexRS for cost savings.
AnswerC

Preemptible worker nodes cut compute costs substantially for fault-tolerant batch ETL, while autoscaling adds or removes workers based on YARN load, matching capacity to the 10x volume growth without over-provisioning. Dataproc handles preemption by re-running lost tasks, so the pipeline stays reliable while satisfying the cost-performance balance.

Why this answer

Preemptible worker nodes significantly reduce cost (up to 80% discount) while autoscaling dynamically adjusts cluster size to match the growing workload, ensuring performance without over-provisioning. This combination handles the 10x data growth efficiently by scaling out during peak loads and scaling in during lulls, using preemptible instances for fault-tolerant tasks like transformation.

Exam trap

The trap here is that candidates often choose Dataflow (Option D) assuming it is always the best for cost and performance, but the question specifically involves Dataproc and batch ETL from Cloud Storage to BigQuery, where preemptible nodes with autoscaling provide a more direct and cost-effective solution without requiring a pipeline rewrite.

How to eliminate wrong answers

Option A is wrong because a single large persistent cluster incurs high costs even when idle, and cannot efficiently handle a 10x growth without manual resizing, leading to either underutilization or performance bottlenecks. Option B is wrong because Cloud Data Fusion is a visual design tool that adds complexity and cost (via Dataproc provisioning) without inherent autoscaling or preemptible node benefits, and is not optimized for batch ETL cost control. Option D is wrong because Dataflow with flexRS is designed for batch workloads with flexible scheduling, but it requires rewriting the pipeline in Apache Beam, which adds migration overhead and may not leverage existing Dataproc investments; flexRS offers cost savings but with potential execution delays, making it less balanced for immediate performance needs.

668
MCQeasy

Which Google Cloud service provides a fully managed, serverless Spark environment without requiring cluster provisioning?

A.Dataproc on GKE
B.Dataflow
C.Dataproc Serverless
D.Cloud Data Fusion
AnswerC

Dataproc Serverless runs Spark workloads without any cluster provisioning, directly satisfying the stem's serverless requirement. Unlike standard Dataproc, which needs manual cluster creation and sizing, it provisions ephemeral compute automatically per job, so no infrastructure management is needed. This makes it the fully managed Spark environment the question describes.

Why this answer

Dataproc Serverless is a fully managed, serverless Spark environment on Google Cloud that eliminates the need to provision or manage clusters. It automatically scales resources and charges only for the duration of the workload, making it ideal for running Spark jobs without infrastructure overhead.

Exam trap

PDE often tests the distinction between serverless and managed services, and candidates may confuse Dataflow (Beam) with Dataproc Serverless (Spark) or think Dataproc on GKE is serverless when it still requires cluster management.

How to eliminate wrong answers

Option A is wrong because Dataproc on GKE requires managing a Kubernetes cluster, which involves cluster provisioning and configuration, not serverless. Option B is wrong because Dataflow is a fully managed service for Apache Beam, not Spark; it is designed for stream and batch processing but does not run Spark jobs. Option D is wrong because Cloud Data Fusion is a fully managed data integration service for building ETL pipelines, but it is not a serverless Spark environment; it may use Dataproc clusters under the hood but requires provisioning.

669
MCQhard

You are designing the deployment process for a Dataflow streaming pipeline that processes financial transactions. The pipeline must be updated without losing in-flight state, such as open windows and timers, and without downtime. Your team uses the Apache Beam Java SDK and deploys from a CI/CD pipeline. Which update strategy should you use?

A.Submit an update with the --update option and a compatible transform change to replace the pipeline in place.
B.Cancel the pipeline and redeploy with a new job name.
C.Use a snapshot and then start a new job from the snapshot to preserve state.
D.Stop the pipeline with a drain, then start a new pipeline from the same template.
AnswerA

Dataflow supports in-place updates via the update option, which preserves pipeline state including windows and timers if the transform changes are compatible. This allows the pipeline to continue processing without downtime. It is the standard mechanism for evolving streaming pipelines while maintaining exactly-once semantics and state.

Why this answer

Dataflow update replaces a running streaming pipeline in place while preserving state, provided the transform changes are compatible with the update compatibility rules. This avoids downtime and keeps windowed state intact. Draining, snapshot-based restart, and cancel-redeploy all interrupt processing or discard state, making them unsuitable for stateful, zero-downtime updates.

Exam trap

The trap here is assuming a snapshot or drain preserves state for a new job, when only an in-place update maintains state without interrupting processing.

670
Multi-Selecthard

A company wants to implement a robust MLOps lifecycle on Google Cloud. Which THREE components are essential?

Select 3 answers
A.Vertex AI Model Registry for versioning
B.Vertex AI Pipelines for orchestration
C.Pub/Sub for event-driven retraining
D.Cloud Build for CI/CD
E.Cloud SQL for model metadata
AnswersA, B, D

Vertex AI Model Registry provides centralised versioning, lineage and stage tracking for trained models. It satisfies the lifecycle constraint by giving every model artifact an immutable, addressable version that pipelines and endpoints can reference, enabling rollback and auditability across retraining cycles.

Why this answer

Vertex AI Model Registry (A) is essential because it provides centralized versioning, lineage, and lifecycle tracking of trained models, which is a core requirement for a robust MLOps lifecycle. Vertex AI Pipelines (B) is essential for orchestrating the end-to-end ML workflow (data prep, training, evaluation, deployment) as reproducible, automated pipeline steps. Cloud Build (D) is essential for CI/CD, enabling automated building, testing, and deployment of ML artifacts and pipeline components as part of the MLOps process.

Pub/Sub (C) can support event-driven retraining but is an optional architectural pattern rather than a core MLOps component, and Cloud SQL (E) is a general relational database not purpose-built for model metadata management, so neither is essential here.

Exam trap

The trap here is that candidates may confuse optional supporting services (like Pub/Sub for event triggers or Cloud SQL for metadata) with the essential components required for a robust MLOps lifecycle, which are versioning, orchestration, and CI/CD.

671
MCQeasy

You need to process a large Spark ML training job on a Dataproc cluster. The job is fault-tolerant and can handle occasional node failures. To reduce costs, which type of worker nodes should you use?

A.Preemptible worker nodes
B.Standard worker nodes
C.High-memory worker nodes
D.Sole-tenant nodes
AnswerA

Preemptible workers cost far less than standard VMs but can be reclaimed at any time. Because the job tolerates occasional node loss, Dataproc simply reschedules the lost tasks, so the fault-tolerance constraint is met while compute spend drops substantially.

Why this answer

Option A is correct because preemptible worker nodes in Dataproc are significantly cheaper than standard VMs but can be reclaimed by Compute Engine at any time. Since the Spark ML job is fault-tolerant and can handle occasional node failures, preemptible workers are the ideal cost-saving choice. Dataproc automatically replaces preempted workers, and Spark's lineage-based recomputation allows the job to continue without manual intervention.

Exam trap

PDE often tests the assumption that preemptible nodes are unsuitable for any production job — candidates miss that fault-tolerant, stateless workloads are exactly the intended use case.

How to eliminate wrong answers

Option B is wrong because standard worker nodes are billed at full on-demand rates and provide no cost advantage for a fault-tolerant workload. Option C is wrong because high-memory nodes cost more than standard nodes and are only justified for memory-intensive workloads, not for general cost reduction. Option D is wrong because sole-tenant nodes are dedicated physical hosts used for compliance/licensing isolation and are far more expensive than preemptible or standard nodes.

672
MCQmedium

A data engineer is using Apache Spark on Dataproc to process a large dataset. They need to perform complex aggregation and transformation with high performance. The dataset has a known schema and they want to take advantage of Catalyst optimizer. Which Spark API should they use?

A.Spark SQL only
B.DataFrames
C.Datasets
D.RDDs
AnswerB

DataFrames expose a structured schema, letting Spark's Catalyst optimizer apply rule-based and cost-based query planning to aggregations and transformations. This delivers the high performance the scenario demands, unlike RDDs, which lack schema awareness and bypass Catalyst entirely.

Why this answer

DataFrames are the correct choice because they combine a declarative, structured API with the Catalyst optimizer, which builds and optimizes a logical plan before execution. With a known schema, DataFrames let Spark apply rule-based and cost-based optimizations such as predicate pushdown, column pruning, and join reordering. This delivers high performance for complex aggregations and transformations without sacrificing the structured API benefits.

Exam trap

The trap here is confusing Spark SQL (a query interface) with DataFrames (the structured API that actually invokes Catalyst), or assuming Datasets are always the best choice when language support and schema handling matter.

How to eliminate wrong answers

Option A is wrong because Spark SQL is a subset/interface for SQL queries and does not by itself provide the full programmatic DataFrame API needed for complex transformations; it is not the broadest answer. Option C is wrong because Datasets are a typed extension available primarily in Scala/Java and are not the general-purpose API for all languages; the question asks for the API that leverages Catalyst with a known schema, which DataFrames do. Option D is wrong because RDDs are the low-level, unstructured API that bypasses Catalyst and Tungsten optimizations, so they cannot take advantage of the optimizer.

673
MCQeasy

Which BigQuery feature allows you to read data directly from Cloud Storage without loading it into BigQuery storage?

A.External tables
B.BI Engine
C.Federated queries
D.Authorized views
AnswerA

External tables define a schema over data residing in Cloud Storage, letting BigQuery query it directly without ingestion. This satisfies the constraint of reading Cloud Storage data without loading it into BigQuery's native storage, unlike native tables which require loading.

Why this answer

External tables in BigQuery allow you to query data directly from Cloud Storage (e.g., CSV, JSON, Avro, Parquet) without loading it into BigQuery's native storage. This is achieved by creating a table that references the external data source, enabling federated queries and on-demand access. The data remains in Cloud Storage, and BigQuery reads it at query time, which is ideal for one-time analysis or when data is frequently updated externally.

Exam trap

PDE often tests the distinction between external tables and federated queries, as both involve querying external data; candidates may confuse the two, but external tables specifically target Cloud Storage, while federated queries target other databases.

How to eliminate wrong answers

Option B is wrong because BI Engine is an in-memory analysis service that accelerates queries on BigQuery data, not a feature for reading external data. Option C is wrong because federated queries typically refer to querying data in external databases (e.g., Cloud SQL) via BigQuery, not directly from Cloud Storage; the term is broader and not specific to Cloud Storage. Option D is wrong because authorized views are used to share query results with specific users or groups, not to read data from external sources.

674
MCQeasy

A company is ingesting real-time sensor data from thousands of devices into Cloud Pub/Sub. They need to process this data with low latency (seconds) and exactly-once semantics. Which data processing service should they use?

A.Cloud Run with Pub/Sub push
B.Cloud Functions triggered by Pub/Sub
C.Dataflow streaming with exactly-once processing
D.Dataproc with Spark Streaming
AnswerC

Dataflow streaming provides the low-latency, per-record processing Pub/Sub ingestion requires, and its exactly-once mode deduplicates via Pub/Sub message IDs and transactional state writes. This satisfies both the seconds-level latency and exactly-once semantics constraints without custom checkpointing.

Why this answer

Dataflow streaming with exactly-once processing is the correct choice because it provides exactly-once semantics for Pub/Sub sources via checkpointing and idempotent sinks, and it meets the low-latency (seconds) requirement through its streaming engine that minimizes per-element overhead. Cloud Dataflow's integration with Pub/Sub ensures that each message is processed exactly once, even in the presence of failures, by using snapshots and consistent state management.

Exam trap

Google Cloud often tests the misconception that serverless services like Cloud Functions or Cloud Run inherently provide exactly-once processing, when in fact they rely on Pub/Sub's at-least-once delivery and require additional logic to achieve exactly-once semantics.

How to eliminate wrong answers

Option A is wrong because Cloud Run with Pub/Sub push does not guarantee exactly-once processing; Pub/Sub push delivery is at-least-once, and Cloud Run's stateless containers cannot enforce exactly-once semantics without external coordination. Option B is wrong because Cloud Functions triggered by Pub/Sub also uses at-least-once delivery from Pub/Sub and lacks built-in mechanisms for exactly-once processing; it is designed for lightweight, event-driven tasks, not for stateful streaming with exactly-once guarantees. Option D is wrong because Dataproc with Spark Streaming provides at-least-once or exactly-once semantics only with additional configuration (e.g., checkpointing and idempotent sinks), but it introduces higher latency (typically seconds to minutes) due to micro-batching and is not optimized for sub-second or low-latency streaming compared to Dataflow's streaming engine.

675
MCQmedium

A data engineer needs to create a Dataflow pipeline template that can be reused across multiple environments (dev, staging, prod) with different parameters (e.g., input Pub/Sub topic, output BigQuery table). Which template type should they use?

A.Dataflow Prime
B.Flex Template
C.Classic Template
D.Cloud Composer workflow template
AnswerB

Flex Templates package the pipeline as a Docker image with runtime parameters supplied via metadata, so the same template runs across dev, staging and prod by passing different Pub/Sub topics and BigQuery tables. Classic templates require recompilation for such changes.

Why this answer

Flex Templates (B) are the correct choice because they package a Dataflow pipeline as a Docker image, allowing environment-specific parameters (e.g., Pub/Sub topic, BigQuery table) to be passed at runtime via the --parameters flag. This enables true reusability across dev, staging, and prod without modifying the template code, unlike Classic Templates which require compile-time parameterization.

Exam trap

The Google Cloud exam often tests the distinction between Classic Templates (compile-time parameterization) and Flex Templates (runtime parameterization), trapping candidates who assume all templates support the same level of parameter flexibility.

How to eliminate wrong answers

Option A is wrong because Dataflow Prime is a managed service for optimizing resource utilization and autoscaling, not a template type for parameterized reuse. Option C is wrong because Classic Templates require parameters to be baked in at staging time, making them less flexible for multi-environment reuse without rebuilding the template. Option D is wrong because Cloud Composer is an Apache Airflow orchestration service used to schedule and monitor workflows, not a Dataflow template type for parameterized pipeline reuse.

Page 8

Page 9 of 10

Page 10

All pages