Courseiva

Google Professional Data Engineer (PDE) — Questions 151–225

747 questions total · 10pages · All types, answers revealed

Page 2

Page 3 of 10

Page 4
151
MCQmedium

A data pipeline uses Cloud Composer (Airflow) to orchestrate Dataproc jobs. Each job submits a Spark application that reads from BigQuery and writes to Cloud Storage. The pipeline runs nightly and takes 6 hours. Management wants to reduce costs. Which approach is most effective?

A.Use preemptible VMs for the Dataproc cluster
B.Switch to Cloud Dataproc billing per second instead of per minute
C.Increase the memory of the driver node to improve performance
D.Upgrade the Cloud Storage class from Standard to Nearline
AnswerA

Dataproc preemptible VMs cost substantially less than standard instances, and Spark's resilient distributed datasets tolerate their eviction by recomputing lost partitions. For a nightly six-hour batch job, that discount directly cuts spend without changing the pipeline's logic.

Why this answer

Preemptible VMs are significantly cheaper (up to 80% discount) than standard VMs and are ideal for fault-tolerant, batch workloads like nightly Dataproc jobs. Since the pipeline runs nightly and takes 6 hours, it can tolerate the occasional preemption of worker nodes by using Spark's built-in resilience (e.g., task retries). This directly reduces compute cost without sacrificing completion, assuming the cluster is configured with enough preemptible workers to handle the workload.

Exam trap

Google Cloud often tests the misconception that 'upgrading' storage class or changing billing granularity saves money, when in fact the correct answer involves leveraging cheaper compute resources (preemptible VMs) that are designed for fault-tolerant batch jobs.

How to eliminate wrong answers

Option B is wrong because Dataproc already bills per second after a 1-minute minimum, so switching to per-second billing is not a change that reduces costs further. Option C is wrong because increasing driver memory does not reduce costs; it may actually increase costs by requiring a larger, more expensive VM, and performance gains are unlikely if the bottleneck is not driver memory. Option D is wrong because upgrading from Standard to Nearline storage increases cost (Nearline has higher retrieval and minimum storage duration fees) and is intended for infrequently accessed data, not for nightly write workloads where data is read soon after writing.

152
MCQhard

A company runs a Dataflow streaming pipeline that processes financial transactions. They need to apply a new transformation that enriches the data with a lookup from Cloud Bigtable without stopping the pipeline. The pipeline must be updated in a way that minimises data loss and preserves exactly-once semantics. What is the recommended approach?

A.Use the Dataflow update option with the same pipeline name and new version, ensuring the transform is backward compatible.
B.Drain the pipeline first, then start a new pipeline with the updated code.
C.Create a new pipeline in parallel and switch the Pub/Sub subscription to the new pipeline.
D.Stop the pipeline, update the code, and restart with a new pipeline name.
AnswerA

Dataflow's update option replaces the pipeline definition while retaining the same name and state, so the Bigtable enrichment transform is applied without draining the pipeline, preserving exactly-once semantics provided the new transform stays backward compatible.

Why this answer

Dataflow supports in-place pipeline updates via the update option, which preserves the pipeline's state (including watermarks and deduplication state) and maintains exactly-once semantics. The new transform must be backward compatible with the existing pipeline's state and schema so that the update can be applied without draining.

Exam trap

PDE often tests the difference between update, drain, and stop — the trap is choosing drain or a parallel pipeline when the question requires preserving exactly-once semantics and minimizing data loss.

How to eliminate wrong answers

Option B is wrong because draining stops ingestion and waits for in-flight data to finish, causing downtime and potentially losing data that arrives during the drain window. Option C is wrong because running a parallel pipeline and switching subscriptions risks duplicate processing and breaks exactly-once guarantees across the cutover. Option D is wrong because stopping and restarting with a new pipeline name discards state, loses exactly-once semantics, and may drop or duplicate in-flight data.

153
MCQeasy

A healthcare analytics team needs to run a series of SQL transformations on data stored in BigQuery. The transformations must run on a schedule, and the team wants to minimize operational overhead by using a fully managed service that integrates with BigQuery and Cloud Logging. They also need to parameterize the SQL queries with runtime values such as the current date. Which Google Cloud service should they use?

A.Cloud Composer with a DAG that uses BigQueryInsertJobOperator to run the SQL queries.
B.BigQuery scheduled queries, using the @run_date parameter for runtime values.
C.Dataflow with a pipeline that reads from BigQuery, applies SQL transformations using Beam SQL, and writes back to BigQuery.
D.Cloud Scheduler triggering a Cloud Function that calls the BigQuery API to execute the SQL.
AnswerB

BigQuery scheduled queries are a fully managed feature that runs SQL on a schedule, supports parameterization with @run_date and other system variables, and integrates with Cloud Logging for monitoring. It requires no infrastructure management, making it ideal for scheduled SQL transformations. This directly meets the team's requirements with minimal operational overhead.

Why this answer

BigQuery scheduled queries are a native, fully managed feature that executes SQL on a defined schedule. They support parameterization with system variables like @run_date, which allows dynamic date-based filtering. Integration with Cloud Logging provides visibility.

This eliminates the need to manage infrastructure or write code, perfectly matching the requirement for a low-overhead, scheduled SQL transformation solution.

Exam trap

The trap here is overcomplicating a simple scheduled SQL task by choosing a general-purpose orchestrator or compute service instead of the native BigQuery scheduling feature.

154
Multi-Selecthard

A Dataflow batch job frequently fails with 'OutOfMemoryError'. Which THREE are common causes? (Choose 3)

Select 3 answers
A.Too many parallel workers
B.Inefficient GroupByKey with hot keys
C.Too many side inputs
D.Too large window accumulation in streaming mode
E.Using Dataflow Shuffle
AnswersB, C, D

GroupByKey materialises all values for each key in a worker's memory before emitting. A hot key concentrates a disproportionate share of records onto one worker, exhausting its heap and triggering OutOfMemoryError despite adequate cluster-wide capacity.

Why this answer

Option B is correct because an inefficient GroupByKey with hot keys forces a single worker to hold all values for a given key in memory during shuffling, causing OutOfMemoryError; using Combine.perKey or pre-aggregation mitigates this. Option C is correct because each side input is loaded into memory on every worker, so too many (or too large) side inputs exhaust worker memory. Option D is correct because in streaming mode, large window accumulation retains all elements for the window in memory until the window fires, and oversized or long windows can exceed worker memory.

Option A is not a cause—more parallel workers distribute load and reduce per-worker memory pressure. Option E is not a cause—Dataflow Shuffle offloads shuffle data to a remote service, reducing worker memory usage rather than causing OOM.

Exam trap

Google Cloud often tests the misconception that increasing parallelism (Option A) always reduces memory errors, but in Dataflow, hot keys cause memory issues regardless of worker count because the hot key's data is processed by a single worker.

155
MCQmedium

Refer to the exhibit. A Dataflow pipeline is failing intermittently with the shown error. Which step should the team take to ensure data quality and prevent such errors?

A.Increase the number of workers to process the data faster.
B.Add a monitoring alert on the 'system_lag' metric.
C.Use a strongly typed schema for the PCollection and let Beam automatically reject malformed data.
D.Modify the pipeline to handle parsing failures by sending invalid records to a dead letter queue.
AnswerD

Routing unparseable records to a dead letter queue isolates malformed data so the pipeline continues processing valid rows, preventing the intermittent failures caused by parsing errors while preserving invalid records for later inspection and remediation.

Why this answer

The error indicates that the pipeline is failing due to malformed or unparseable data. By sending invalid records to a dead letter queue (DLQ), the pipeline can continue processing valid data while capturing and isolating bad records for later analysis or reprocessing. This pattern is a standard data quality practice in Apache Beam and Dataflow, ensuring that transient or corrupt data does not cause pipeline failures.

Exam trap

Google Cloud often tests the distinction between scaling solutions (like increasing workers) and data quality patterns (like dead letter queues), trapping candidates who confuse performance optimization with error handling.

How to eliminate wrong answers

Option A is wrong because increasing the number of workers addresses throughput and latency, not data quality or malformed data errors; it does not prevent parsing failures. Option B is wrong because monitoring the 'system_lag' metric tracks pipeline latency, not data quality issues; it would not prevent or handle malformed records. Option C is wrong because while strongly typed schemas can help catch type mismatches at compile time, they do not automatically reject malformed data at runtime in Beam; the pipeline would still fail if a record cannot be parsed into the schema, and Beam does not have built-in automatic rejection to a dead letter queue without explicit handling.

156
MCQmedium

A company deploys a machine learning model to Vertex AI for real-time predictions. After deployment, they notice that prediction latency spikes during peak traffic hours. Which approach should they take to reduce latency without sacrificing accuracy?

A.Configure auto-scaling with higher min and max instances
B.Reduce the number of input features
C.Switch from online to batch prediction
D.Use a larger machine type for the model
AnswerA

Auto-scaling with higher minimum and maximum instances adds serving capacity during peak traffic, absorbing load and cutting latency while the same model preserves accuracy. This satisfies the stem's constraint of reducing latency without sacrificing accuracy.

Why this answer

Configuring auto-scaling with higher min and max instances ensures that Vertex AI has sufficient pre-warmed replicas to handle traffic spikes without cold-start latency. This approach maintains model accuracy because it does not alter the model architecture or inference logic, only the infrastructure capacity.

Exam trap

Google Cloud often tests the misconception that reducing features or using batch prediction is the primary way to reduce latency, but the real exam trap is that candidates overlook the need to maintain real-time capability and accuracy, and instead choose a solution that changes the model or prediction mode rather than scaling infrastructure.

How to eliminate wrong answers

Option B is wrong because reducing the number of input features may degrade model accuracy, and the question explicitly requires not sacrificing accuracy. Option C is wrong because switching from online to batch prediction eliminates real-time capability, which contradicts the requirement for real-time predictions. Option D is wrong because using a larger machine type can reduce latency but often increases cost and may introduce cold-start delays if scaling is not addressed; it does not directly solve latency spikes during peak traffic, and the question asks for a solution that does not sacrifice accuracy, which a larger machine type does not affect but is not the most targeted fix for traffic-induced latency.

157
MCQhard

You are migrating an on-premises PostgreSQL database to Cloud SQL. You need to continuously replicate changes to BigQuery for real-time analytics with minimal latency. Which service should you use?

A.Dataflow with JDBC source
B.Pub/Sub with a Cloud Function that writes to BigQuery
C.Storage Transfer Service
D.Datastream
AnswerD

Datastream provides serverless change data capture, streaming PostgreSQL changes continuously into BigQuery with minimal latency. Unlike batch extracts or scheduled jobs, it replicates ongoing changes in near real time, meeting the continuous replication and low-latency analytics constraint.

Why this answer

Datastream is Google Cloud's serverless change data capture (CDC) service that continuously replicates changes from sources like PostgreSQL, MySQL, Oracle, and SQL Server into BigQuery or Cloud Storage with low latency. It reads the source's replication log and streams changes without impacting production workloads significantly. This directly matches the requirement for continuous, near-real-time replication into BigQuery.

Exam trap

PDE often tests CDC versus batch ingestion; candidates pick Dataflow because it is the general-purpose pipeline tool, but Dataflow alone does not provide native database CDC.

How to eliminate wrong answers

Option A is wrong because Dataflow with a JDBC source performs batch or scheduled reads, not continuous CDC, so it cannot deliver minimal-latency change replication. Option B is wrong because Pub/Sub with a Cloud Function requires custom application logic to capture and publish changes, and it does not natively read PostgreSQL replication logs. Option C is wrong because Storage Transfer Service moves files between object stores and has no database CDC capability.

158
MCQeasy

A data scientist has trained an XGBoost model on Vertex AI and wants to deploy it to an endpoint with automatic scaling based on traffic. What is the recommended deployment approach?

A.Export the model to a container and deploy on Cloud Run
B.Use AI Platform Prediction with batch prediction
C.Deploy the model as an API on App Engine
D.Use Vertex AI Endpoints with automatic scaling enabled
AnswerD

Vertex AI Endpoints natively support automatic scaling, adjusting replica counts to match traffic without manual intervention. Deploying the trained XGBoost model there satisfies the stem's traffic-based scaling requirement, unlike batch prediction or a fixed single-replica deployment.

Why this answer

Vertex AI Endpoints with automatic scaling enabled is the recommended approach because it directly supports deploying trained models (including XGBoost) as online prediction endpoints with built-in autoscaling based on incoming traffic. This service manages the underlying infrastructure, load balancing, and scaling policies, aligning with the requirement for automatic scaling without additional containerization or serverless overhead.

Exam trap

Google Cloud often tests the distinction between online (real-time) and batch prediction services, and the trap here is that candidates may confuse Vertex AI Endpoints with generic serverless options like Cloud Run or App Engine, overlooking the fact that Vertex AI provides a purpose-built, managed endpoint service with native autoscaling for ML models.

How to eliminate wrong answers

Option A is wrong because exporting the model to a container and deploying on Cloud Run requires manual containerization and does not natively integrate with Vertex AI's model registry, versioning, or monitoring, and Cloud Run's scaling is based on request concurrency rather than the model-specific metrics Vertex AI provides. Option B is wrong because AI Platform Prediction with batch prediction is designed for offline, asynchronous predictions on large datasets, not for real-time online serving with automatic scaling based on live traffic. Option C is wrong because deploying the model as an API on App Engine introduces unnecessary complexity and lacks the optimized serving infrastructure, model versioning, and traffic splitting capabilities that Vertex AI Endpoints offer for ML models.

159
MCQmedium

A team wants to use Cloud Pub/Sub Lite for a high-throughput, low-cost messaging system. They need exactly-once delivery to subscribers. What should they know about Pub/Sub Lite's delivery guarantees?

A.Pub/Sub Lite provides at-least-once delivery, same as standard Pub/Sub.
B.Pub/Sub Lite provides exactly-once delivery when using push subscriptions.
C.Pub/Sub Lite provides exactly-once delivery when using pull subscriptions.
D.Pub/Sub Lite supports exactly-once delivery by default.
AnswerA

Pub/Sub Lite guarantees at-least-once delivery, so duplicate messages can reach subscribers; exactly-once is unavailable. This satisfies the stem's constraint by clarifying that the team cannot rely on deduplication, and must implement idempotent processing or their own deduplication logic to achieve exactly-once semantics.

Why this answer

Pub/Sub Lite offers at-least-once delivery like standard Pub/Sub; exactly-once is not guaranteed.

160
MCQhard

A company has a model that requires GPU for inference and has strict latency requirements. They deployed on Vertex AI Endpoint with autoscaling but observe cold start latency when scaling up. What is the best solution?

A.Set a higher min_replica_count to keep instances warm
B.Pre-compile the model with TensorRT
C.Use a larger GPU instance
D.Switch to batch prediction
AnswerA

Setting a higher min_replica_count keeps a baseline of GPU-backed instances running continuously, so scaling events add capacity without provisioning new nodes. This eliminates the cold start latency that violates the strict latency requirement, since requests never wait for instance initialisation.

Why this answer

Setting a higher min_replica_count ensures that a baseline number of GPU instances are always running and ready to serve inference requests, eliminating cold start latency because new instances do not need to be provisioned and loaded from scratch when traffic spikes. This directly addresses the autoscaling-induced cold start issue by maintaining a warm pool of replicas.

Exam trap

The trap here is that candidates often confuse inference optimization techniques (like TensorRT or larger GPUs) with infrastructure-level scaling configurations, failing to recognize that cold start is a provisioning delay, not a compute performance issue.

How to eliminate wrong answers

Option B is wrong because pre-compiling the model with TensorRT optimizes inference performance (e.g., reducing latency per request) but does not eliminate the cold start latency that occurs when new instances are spun up from zero. Option C is wrong because using a larger GPU instance reduces per-request compute time but does not prevent the provisioning and model-loading delay when scaling from zero replicas. Option D is wrong because switching to batch prediction is designed for asynchronous, non-real-time workloads and does not meet strict latency requirements; it also does not address cold start for online inference.

161
MCQeasy

A media company stores final video masters in a Cloud Storage bucket. Regulatory rules require that each object be unalterable for seven years, and the company must be able to prove retention compliance to auditors. Objects are written once and never edited. The data engineer needs the strongest native Cloud Storage control that prevents deletion or overwrite for the required period. What should the engineer configure?

A.A bucket retention policy with a seven-year retention period and a locked retention policy.
B.Uniform bucket-level access with IAM roles limited to a single compliance group.
C.Object Versioning on the bucket plus a lifecycle rule to delete noncurrent versions.
D.A signed URL with a seven-year expiration that is shared only with the compliance team.
AnswerA

A bucket retention policy prevents deletion or replacement of objects until the retention period elapses, and locking the policy makes it permanent so it cannot be shortened or removed. This gives the unalterable, provable seven-year guarantee the auditors require, and it is the native Cloud Storage control designed for exactly this write-once retention need.

Why this answer

A locked bucket retention policy is the native Cloud Storage feature that enforces immutability for a fixed period and cannot be weakened once locked. It directly satisfies the write-once, provable-retention requirement, whereas versioning, signed URLs, and IAM policies only govern recovery or access. For regulatory retention of final masters, the locked retention policy is the correct control.

Exam trap

The trap here is confusing access control or versioning with true immutability, when only a locked retention policy legally prevents deletion.

162
MCQeasy

A company is building a data lake on Cloud Storage for log analysis. Log files (CSV) arrive every 5 minutes from multiple sources. The files should be ingested into BigQuery for reporting within 15 minutes. Which approach best meets the requirements with minimal operational overhead?

A.Set up a Cloud Storage notification to trigger a Cloud Function that loads each file into BigQuery using the BigQuery API.
B.Schedule a daily batch load from Cloud Storage to BigQuery using the BigQuery Data Transfer Service.
C.Use Dataflow to read from Pub/Sub (ingested from Cloud Storage) and write to BigQuery.
D.Use BigQuery federated queries to query the CSV files directly from Cloud Storage.
AnswerA

Cloud Storage notifications triggering a Cloud Function gives event-driven, per-file ingestion, so each CSV lands in BigQuery within the 15-minute window. Serverless execution removes cluster or pipeline management, satisfying the minimal operational overhead constraint. The BigQuery API load job handles schema and append semantics directly, avoiding scheduled batch polling that could breach the latency requirement.

Why this answer

Cloud Storage notifications trigger a Cloud Function on each file upload, which then loads the file into BigQuery via the BigQuery API. This provides near-real-time ingestion (within seconds of file arrival) with minimal operational overhead, as there are no servers to manage and no scheduling needed. The 5-minute file arrival and 15-minute SLA are easily met without complex infrastructure.

Exam trap

Google Cloud often tests the misconception that serverless options like Cloud Functions are only for simple tasks, but here they are the most efficient choice for near-real-time ingestion with minimal overhead, while Dataflow is overkill for this straightforward file-load pattern.

How to eliminate wrong answers

Option B is wrong because a daily batch load does not meet the 15-minute ingestion requirement; it would only load data once per day, causing up to 24 hours of latency. Option C is wrong because it introduces unnecessary complexity and operational overhead by adding Pub/Sub and Dataflow, which are not needed when files are already in Cloud Storage and can be loaded directly via a Cloud Function. Option D is wrong because BigQuery federated queries do not ingest data into BigQuery; they query the CSV files directly from Cloud Storage, which is slower and does not support the required reporting use case where data must be stored in BigQuery for efficient analysis.

163
Drag & Dropmedium

Drag and drop the steps to deploy a Cloud Dataflow pipeline from a template into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Deploying a Dataflow pipeline from a template via the Cloud Console involves a specific sequence: first, navigate to the Dataflow page, then create a job from a template, select the appropriate template, fill in the required parameters (such as input/output and temp location), and finally run the job. This order ensures all necessary configurations are applied correctly before execution.

164
MCQmedium

A data engineer needs to query data from BigQuery and another cloud provider's storage (AWS S3) using a single SQL query. The data must not be moved or copied to GCP. Which Google Cloud service should they use?

A.Cloud Storage Transfer Service
B.Dataplex
C.BigQuery Data Transfer Service
D.BigQuery Omni
AnswerD

BigQuery Omni runs BigQuery SQL directly against data held in AWS S3 through Anthos-hosted compute, returning results without extracting or copying the data into Google Cloud. This satisfies the constraint that the S3 data must remain in place while a single query spans both sources.

Why this answer

BigQuery Omni allows querying data across multiple clouds (AWS S3, Azure Blob Storage) using BigQuery's interface without moving data. BigQuery Omni runs compute in the other cloud's region. BigQuery Transfer Service moves data into BigQuery.

Dataplex is for data management, not cross-cloud queries. Cloud Storage Transfer Service is for moving data between clouds.

165
MCQmedium

A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?

A.Use a single-node cluster to eliminate overhead
B.Enable high-availability mode to avoid restarts
C.Use preemptible instances for worker nodes
D.Increase the number of standard workers to finish faster
AnswerC

Preemptible instances cost significantly less than standard Dataproc worker nodes, and Spark can tolerate their loss through retries and task rescheduling. Since the daily jobs are short and their characteristics stay unchanged, using preemptible workers cuts compute spend without altering the job design.

Why this answer

Preemptible (Spot) VMs cost up to 80% less than standard VMs and are ideal for fault-tolerant, batch-oriented workloads like Spark ML jobs that can tolerate occasional preemption. Since the jobs run only 2 hours daily and Dataproc automatically handles node replacement when preemptible instances are reclaimed, using preemptible workers delivers the largest cost reduction without changing job characteristics.

Exam trap

The trap is assuming 'reduce cost' means 'reduce cluster size' — candidates pick single-node or fewer workers, but the exam expects you to recognize that preemptible/Spot instances reduce cost per unit while preserving job characteristics.

How to eliminate wrong answers

Option A is wrong because a single-node cluster eliminates worker parallelism entirely, which would drastically slow or break distributed Spark ML jobs — it reduces cost by removing capacity, not by optimizing it. Option B is wrong because high-availability mode adds a second master node, increasing cost rather than reducing it, and it addresses master reliability, not worker cost. Option D is wrong because adding more standard workers increases cost linearly and only helps if the job is worker-bound; it does not reduce per-unit cost.

166
MCQhard

An organization wants to enforce that data in a Cloud Storage bucket cannot be deleted or overwritten for 7 years due to regulatory compliance. Which Cloud Storage feature should they use?

A.Retention Policy with Bucket Lock
B.IAM conditions
C.Object Lifecycle Management
D.Object holds
AnswerA

A retention policy combined with Bucket Lock prevents objects from being deleted or overwritten until the retention period expires, enforcing immutability for the full seven years. This satisfies the regulatory requirement, whereas lifecycle rules and versioning permit deletion or alteration.

Why this answer

Retention Policy with a retention period ensures objects cannot be deleted or overwritten during that period. Bucket Lock makes the policy permanent. Object holds are per-object.

Lifecycle management automates transitions/deletions, opposite of retention.

167
MCQeasy

A data engineer needs to design a batch processing pipeline using Cloud Data Fusion. The pipeline should read data from Cloud Storage, perform transformations (join, filter, aggregate), and write to BigQuery. What is the most efficient way to handle the transformations?

A.Use Data Fusion Wrangler to visually design the transformations and then run the pipeline on a Dataproc cluster.
B.Use SQL queries in BigQuery to perform the transformations after loading raw data into staging tables.
C.Use custom Python scripts in a Cloud Function triggered after the files land in Cloud Storage.
D.Use Apache Spark on Dataproc to code the transformations manually, bypassing Data Fusion.
AnswerA

Wrangler's visual directives compile into Spark code executed on the Dataproc cluster provisioned by Cloud Data Fusion, so joins, filters and aggregations run in parallel rather than in the driver. This satisfies the batch pipeline's efficiency constraint, since the same distributed engine handles both the Cloud Storage read and the BigQuery write.

Why this answer

Cloud Data Fusion Wrangler provides a visual, no-code interface for designing transformations (join, filter, aggregate) that are then compiled into an Apache Spark or MapReduce program and executed on a Dataproc cluster. This approach leverages Data Fusion's native integration with Dataproc for efficient, scalable batch processing without manual coding, while keeping the pipeline fully managed within the Data Fusion ecosystem.

Exam trap

Google Cloud often tests the misconception that Cloud Data Fusion is only a visual tool and that transformations must be coded manually in Spark or SQL, when in fact Wrangler generates optimized Spark code under the hood and integrates seamlessly with Dataproc for execution.

How to eliminate wrong answers

Option B is wrong because it bypasses Data Fusion entirely, requiring raw data to be loaded into BigQuery staging tables first, which adds latency and storage costs; transformations in BigQuery are better suited for analytics queries, not as a primary ETL step in a Data Fusion pipeline. Option C is wrong because Cloud Functions have a maximum timeout of 9 minutes (540 seconds) and limited memory (up to 8 GB), making them unsuitable for large-scale batch transformations like joins and aggregations on datasets that may be gigabytes or terabytes in size. Option D is wrong because it suggests manually coding Spark on Dataproc, which defeats the purpose of using Data Fusion's visual design and managed execution; while Spark can be used, Data Fusion already abstracts and optimizes the Spark execution, so manual coding adds unnecessary complexity and maintenance overhead.

168
MCQmedium

Your company uses Pub/Sub to ingest clickstream data. Messages must be processed in order for the same user_id. How should you configure the Pub/Sub subscription to guarantee ordering?

A.Use a pull subscription with enable_message_ordering=true
B.Use a pull subscription with exactly-once delivery enabled
C.Use a push subscription with acknowledgement deadline set to 600 seconds
D.Use a push subscription with a dead letter topic
AnswerA

Ordering in Pub/Sub is enforced per ordering key, so setting enable_message_ordering=true on a pull subscription makes the service deliver messages sharing the same user_id sequentially. Without this flag, Pub/Sub provides at-least-once delivery with no ordering guarantee, so clickstream events for one user could arrive out of sequence.

Why this answer

Pub/Sub message ordering is enabled per-subscription by setting enable_message_ordering=true, and ordering keys (such as user_id) must be set on published messages. With ordering enabled, Pub/Sub delivers messages with the same ordering key in publish order to a single subscriber, satisfying the per-user_id ordering requirement.

Exam trap

The trap is confusing exactly-once delivery with ordering — candidates see 'guarantee' and pick exactly-once, but ordering and exactly-once are orthogonal Pub/Sub features that must be enabled separately.

How to eliminate wrong answers

Option B is wrong because exactly-once delivery guarantees no duplicate delivery but does not guarantee ordering — these are independent features, and exactly-once alone will not preserve per-key order. Option C is wrong because extending the ack deadline to 600 seconds only affects redelivery timing; it has no bearing on message ordering. Option D is wrong because a dead letter topic captures messages that fail processing after retries, which is a reliability pattern, not an ordering mechanism.

169
MCQhard

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle late-arriving data and emit correct results. You need to ensure that the pipeline's windowing and triggering strategy produces accurate aggregations. Which combination of windowing and triggering should you use?

A.Session windows with default trigger and no allowed lateness.
B.Fixed windows with a trigger that fires after watermark passes and allowed lateness greater than 0.
C.Sliding windows with repeated trigger and allowed lateness set to a positive value.
D.Fixed windows with default trigger and allowed lateness of 0.
AnswerB

Fixed windows with a trigger that fires after the watermark passes the window end, combined with allowed lateness greater than 0, ensures that late data within the allowed lateness period is included in the window's aggregation. This setup produces accurate results for both on-time and late-arriving data. The trigger can also be configured to fire early or on repeated updates, but the key is to allow lateness so that late elements are not dropped.

Why this answer

To handle late-arriving data and produce accurate aggregations, you need a windowing strategy that allows late elements to be incorporated. Fixed windows with a trigger that fires after the watermark and a positive allowed lateness achieve this: the trigger emits results when the watermark passes, and any late data within the allowed lateness is added to the window, causing the trigger to fire again and update the results. This ensures completeness and correctness.

Exam trap

The trap here is assuming that the default trigger with zero allowed lateness is sufficient for late data, when in fact late elements are discarded after the watermark passes unless allowed lateness is set.

170
MCQmedium

A company wants to use dbt (data build tool) to transform data in BigQuery. They have a Cloud Storage bucket containing raw CSV files that are loaded daily into BigQuery via an external table. Which dbt feature should they use to modularize the transformation logic and handle dependencies between models?

A.dbt tests
B.dbt snapshots
C.dbt models with ref()
D.dbt seeds
AnswerC

The ref() function creates a directed acyclic graph between dbt models, letting each model reference upstream ones by name rather than hard-coded table paths. This satisfies the modularisation and dependency-handling requirement, since dbt resolves build order automatically and materialises each model against the BigQuery external table.

Why this answer

C is correct because dbt models with the `ref()` function allow you to modularize SQL transformation logic and automatically handle dependencies between models. When you use `ref('model_name')`, dbt builds a dependency graph, ensuring models are executed in the correct order based on their references. This is essential for transforming raw data from an external table into a structured, analytics-ready dataset in BigQuery.

Exam trap

Candidates often confuse the purpose of dbt components: models with `ref()` manage transformation logic and dependencies, while tests handle data quality, snapshots track historical changes, and seeds load static data.

How to eliminate wrong answers

Option A is wrong because dbt tests are used for validating data quality (e.g., uniqueness, not null) and do not handle transformation logic or dependency management. Option B is wrong because dbt snapshots are designed to capture historical changes in slowly changing dimensions (Type 2 SCDs), not to modularize transformation logic or manage model dependencies. Option D is wrong because dbt seeds are used to load static CSV files directly into the warehouse as tables, not to transform data or manage dependencies between models.

171
Multi-Selectmedium

A healthcare company needs to ingest HL7 messages from an on-premises system into Google Cloud. The messages arrive continuously and must be processed in near-real-time, with transformations applied before loading into BigQuery. The company wants to use a fully managed service for message ingestion and a serverless data processing service. They also need to ensure that the pipeline can handle bursts of traffic and that the data is encrypted at rest. Which TWO Google Cloud services should they use? (Choose two.)

Select 2 answers
A.Cloud Data Fusion
B.Cloud Dataflow
C.Cloud Dataproc
D.Cloud Pub/Sub
E.Cloud Composer
AnswersB, D

Dataflow is a fully managed, serverless service for stream and batch processing. It can read from Pub/Sub, apply transformations to HL7 messages, and write to BigQuery. It autoscales to handle traffic bursts and integrates with Cloud KMS for encryption. Dataflow provides exactly-once processing when needed and is ideal for near-real-time transformations.

Why this answer

Pub/Sub provides scalable, managed ingestion of HL7 messages with encryption at rest, and Dataflow offers serverless stream processing with autoscaling and exactly-once semantics. Together, they form a fully managed pipeline that can handle bursts and transform data before loading into BigQuery. Other services are either not serverless or not designed for real-time processing.

Exam trap

The trap here is selecting Dataproc or Data Fusion because they are data processing services, but they are not fully serverless and may require cluster management, which contradicts the requirement.

172
MCQhard

A financial services company must retain trade records for seven years in Cloud Storage. Regulators require that no object can be deleted or overwritten before its retention period expires, even by project owners, and that the policy cannot be removed. The company also needs to prove compliance during audits. Which combination of controls should they implement?

A.Enable Object Versioning and apply an Organization Policy that denies storage.objects.delete for all principals.
B.Configure a lifecycle rule that transitions objects to Archive storage after 30 days and deletes them after seven years.
C.Use Customer-Managed Encryption Keys in Cloud KMS and revoke key access after each upload.
D.Set a bucket retention policy and lock it, then upload objects with a retention period covering seven years.
AnswerD

A bucket retention policy prevents deletion or replacement of objects until their retention period expires, and once the policy is locked it cannot be removed or shortened. This provides the WORM behavior regulators require and produces audit evidence. Applying it to the bucket and setting object retention for seven years ensures every trade record is protected for the mandated duration, even from project owners.

Why this answer

A locked bucket retention policy enforces WORM semantics: objects cannot be deleted or overwritten until their retention period elapses, and the lock prevents removal or reduction of the policy. Uploading records with a seven-year retention period satisfies the regulatory timeline and creates auditable proof. Other controls either can be changed, do not block deletion, or address confidentiality rather than immutability.

Exam trap

The trap here is confusing encryption key control or lifecycle rules with immutability, when only a locked retention policy prevents deletion and modification for a fixed period.

173
Multi-Selectmedium

An organization is moving on-premises Hadoop workloads to Google Cloud. They need to minimize code changes and manage transient clusters for cost savings. Which two Google Cloud services should they consider? (Choose TWO.)

Select 2 answers
A.Compute Engine with self-managed Hadoop
B.BigQuery
C.Dataproc on GKE
D.Cloud Dataproc
E.Cloud Dataflow
AnswersC, D

Dataproc on GKE runs Spark and Hadoop workloads inside Kubernetes, so existing jobs need minimal modification. It supports ephemeral, autoscaling clusters that terminate when idle, directly satisfying the transient-cluster cost constraint. This makes it a valid choice alongside standard Dataproc for the migration scenario.

Why this answer

Cloud Dataproc (option D) is a managed service for running Spark and Hadoop clusters. It supports transient clusters that can be created on-demand and deleted when idle, minimizing costs. It also allows direct migration of on-premises Hadoop code with minimal changes because it supports standard Hadoop/Spark APIs.

Dataproc on GKE (option C) provides similar benefits but runs containerized workloads on GKE, offering additional ephemeral cluster capabilities and integration with Kubernetes. Both options minimize code changes and enable transient clusters for cost savings, while BigQuery (option B) requires rewriting SQL queries and Cloud Dataflow (option E) requires converting to Beam pipelines. Compute Engine with self-managed Hadoop (option A) does not provide transient cluster management by default.

Exam trap

The trap here is that candidates often confuse Cloud Dataflow (a Google Cloud service that runs Beam pipelines) with Dataproc, not realizing that Dataflow requires rewriting Hadoop jobs into Beam pipelines, while Dataproc on GKE and Cloud Dataproc directly support unmodified Hadoop/Spark code.

174
Multi-Selecthard

Which TWO strategies help reduce prediction latency for a real-time model deployed on Vertex AI Endpoint?

Select 2 answers
A.Use batch prediction instead of online
B.Use Cloud CDN to cache predictions
C.Use a larger machine type (e.g., n1-highcpu-16)
D.Reduce model complexity (e.g., quantize or prune)
E.Enable autoscaling with a minimum replica count
AnswersD, E

Quantisation and pruning shrink the model's parameter count and computational graph, cutting the arithmetic operations each inference performs. Fewer operations mean faster forward passes, directly lowering per-request prediction latency on the Vertex AI Endpoint. This satisfies the stem's real-time latency constraint without altering the endpoint's autoscaling or traffic configuration.

Why this answer

Option D is correct because reducing model complexity via quantization or pruning shrinks the compute and memory footprint of each inference, directly lowering the time the model takes to produce a prediction on the Vertex AI Endpoint. Option E is correct because enabling autoscaling with a minimum replica count keeps warm instances available so incoming real-time requests avoid cold-start delays and are served immediately, reducing prediction latency. Option A is wrong because batch prediction is an offline, asynchronous mode and does not serve real-time online requests.

Option B is wrong because Cloud CDN caches HTTP content, not model inference results, and Vertex AI Endpoint predictions are not cacheable that way. Option C is wrong because simply choosing a larger machine type does not guarantee lower latency and may not address the actual bottleneck, unlike reducing model complexity or keeping replicas warm.

Exam trap

Google often tests the misconception that scaling up hardware (Option C) is the primary way to reduce latency, when in fact model optimization (Option D) and proper autoscaling (Option E) are more effective and cost-efficient strategies for real-time inference.

175
MCQmedium

A data engineer is designing a BigQuery table for a fraud detection system. The table will store 500 TB of transaction records and will be queried constantly by filtering on a transaction timestamp column and a customer_id column. The engineer needs to minimize the amount of data scanned by these queries. What should the engineer do?

A.Create a materialized view that pre-aggregates transactions by customer_id and timestamp.
B.Use the BigQuery Storage Write API to stream data into the table with a fixed schema.
C.Set the table's expiration time to 30 days to limit the amount of historical data stored.
D.Partition the table by the transaction timestamp column and cluster by the customer_id column.
AnswerD

Partitioning by transaction timestamp restricts scans to the relevant time ranges, and clustering by customer_id further reduces the data read within each partition. This combination is the standard BigQuery optimization for high-volume tables filtered on time and a high-cardinality key, directly minimizing bytes scanned and cost.

Why this answer

Partitioning by transaction timestamp and clustering by customer_id is the recommended approach for large BigQuery tables that are frequently filtered on a time column and a high-cardinality key. Partition pruning limits scanned partitions, and clustering sorts data within partitions so that filters on customer_id read fewer blocks, lowering both cost and latency.

Exam trap

The trap here is assuming that any performance feature, such as a materialized view or the Storage Write API, will reduce bytes scanned, when only partitioning and clustering directly affect query scan cost.

176
MCQmedium

Your organization has a BigQuery flat-rate reservation with 500 slots. During peak hours, queries are queued and you need additional capacity temporarily. You want to add slots for a burst of activity without committing to a long-term purchase. What should you do?

A.Switch to on-demand pricing
B.Use flex slots
C.Create a secondary reservation with autoscaling
D.Purchase additional committed use reservations
AnswerB

Flex slots add short-term BigQuery capacity in 500-slot increments for as little as 60 seconds, then delete automatically. They satisfy the stem's temporary burst requirement without a long-term commitment, unlike annual or three-year commitments, and supplement the existing flat-rate reservation's 500 slots during peak queueing.

Why this answer

Flex slots are a BigQuery feature that allows you to purchase slots for a short, fixed duration (e.g., 1 hour, 1 day, 1 week) without a long-term commitment. They are designed exactly for temporary bursts of query activity, such as peak-hour spikes, and can be added to an existing flat-rate reservation. By using flex slots, you temporarily increase the slot capacity of your reservation, reducing queuing, and then the slots automatically expire after the chosen period, avoiding ongoing costs.

Exam trap

PDE often tests the distinction between temporary and committed capacity options, and candidates may confuse flex slots with autoscaling or committed use discounts, leading them to pick options that imply long-term commitments or automatic scaling rather than a short-term, manual burst.

How to eliminate wrong answers

Option A is wrong because switching to on-demand pricing would remove the flat-rate reservation and its predictable cost model, and on-demand pricing is not a temporary add-on; it's a completely different billing model that may not provide the dedicated capacity needed for the burst. Option C is wrong because creating a secondary reservation with autoscaling is not a temporary solution—autoscaling reservations require a baseline commitment and are meant for sustained, variable workloads, not short bursts; moreover, you cannot simply create a secondary reservation without additional cost commitments. Option D is wrong because purchasing additional committed use reservations involves a long-term commitment (1 or 3 years), which contradicts the requirement to avoid a long-term purchase.

177
MCQeasy

A data scientist has iterated on a model and produced a new version. The organization requires the ability to roll back to the previous version quickly if the new version performs poorly in production. Which approach should be used?

A.Store each model version in a separate Cloud Storage bucket.
B.Keep the previous model in a container image and redeploy via Cloud Run.
C.Use Cloud Source Repositories to tag model versions.
D.Upload both versions to Vertex AI Model Registry and use endpoint traffic splitting to route 100% to the safe version if needed.
AnswerD

Vertex AI Model Registry stores both model versions with lineage, and endpoint traffic splitting routes a percentage of requests to each. Sending 100% to the prior version achieves immediate rollback without redeployment, satisfying the fast-reversion requirement.

Why this answer

Vertex AI Model Registry allows you to deploy multiple model versions and use endpoint traffic splitting to gradually shift traffic or instantly route 100% to a specific version. This enables immediate rollback by setting the traffic split to 100% for the previous model version without redeploying or changing infrastructure.

Exam trap

Google Cloud often tests the misconception that version control tools (like Cloud Source Repositories) or storage buckets are sufficient for rollback, when in fact the key requirement is a managed model registry with traffic splitting capabilities for instant, no-downtime rollback.

How to eliminate wrong answers

Option A is wrong because storing each model version in a separate Cloud Storage bucket does not provide a mechanism for quick rollback; you would still need to redeploy the model from that bucket, which is not instantaneous. Option B is wrong because keeping the previous model in a container image and redeploying via Cloud Run is not a rollback strategy—it requires a new deployment, which takes time and does not leverage Vertex AI's managed traffic splitting. Option C is wrong because Cloud Source Repositories is a source code version control service, not a model registry; tagging model versions there does not affect production endpoint traffic.

178
MCQhard

A data engineer is loading a 2 TB CSV dataset from Cloud Storage into a partitioned BigQuery table. The load job fails with an error indicating too many errors during parsing. The source files have inconsistent column counts and some rows contain embedded newlines. The engineer wants to load the data reliably with minimal preprocessing. What should the engineer do?

A.Load the CSV files into a single external table and query it with the ignoreUnknownValues option.
B.Set the maxBadRecords option to a high value and rerun the load job with the same CSV files.
C.Convert the CSV files to newline-delimited JSON or Parquet, then load the converted files into BigQuery.
D.Increase the load job's maximum parallelism by splitting the CSV files into smaller objects in Cloud Storage.
AnswerC

Newline-delimited JSON and Parquet handle embedded newlines and variable schemas without relying on delimiter-based parsing. Converting the source removes the ambiguity that caused the load failure, and BigQuery loads these formats natively with schema autodetection or an explicit schema, yielding reliable ingestion with minimal manual row fixing.

Why this answer

CSV parsing is delimiter-based and breaks when rows contain embedded newlines or inconsistent column counts. Converting to newline-delimited JSON or Parquet removes delimiter ambiguity and lets BigQuery load the data reliably, addressing the root cause rather than suppressing errors or increasing throughput.

Exam trap

The trap here is treating the load error as a throughput or error-tolerance problem, when the real issue is that CSV parsing cannot represent embedded newlines without careful quoting.

179
MCQeasy

A company wants to run complex analytical queries on structured data without managing infrastructure. The data volume is terabytes and queries can take seconds to minutes. Which service is appropriate?

A.Firestore
B.Cloud Bigtable
C.Cloud SQL
D.BigQuery
AnswerD

BigQuery suits this scenario because its serverless, columnar architecture executes complex analytical SQL across terabyte datasets without infrastructure provisioning, satisfying the no-management constraint. Its distributed query engine returns results in seconds to minutes, matching the stated latency tolerance, unlike transactional row-store databases or single-node engines.

Why this answer

BigQuery is correct because it is a serverless, highly scalable data warehouse designed for running complex analytical queries on terabytes of data with fast query performance (seconds to minutes) without any infrastructure management. It uses a columnar storage format and a distributed query engine to handle large-scale structured data efficiently, making it ideal for this use case.

Exam trap

Google Cloud exams often test the distinction between OLTP databases (Cloud SQL, Firestore, Bigtable) and OLAP/data warehouse services (BigQuery), where candidates mistakenly choose Cloud SQL for analytical workloads due to its SQL familiarity, ignoring its scalability and performance limitations for large-scale analytics.

How to eliminate wrong answers

Option A is wrong because Firestore is a NoSQL document database optimized for real-time mobile and web app data with low-latency reads/writes, not for complex analytical queries on terabytes of data. Option B is wrong because Cloud Bigtable is a wide-column NoSQL database designed for high-throughput, low-latency operational workloads (e.g., time-series, IoT), not for complex analytical queries that require SQL-like joins and aggregations. Option C is wrong because Cloud SQL is a managed relational database for OLTP workloads with limited scalability (up to ~30 TB), and running complex analytical queries on terabytes of data would cause performance bottlenecks and require manual sharding or read replicas.

180
MCQeasy

You are operating a streaming data pipeline that uses Cloud Pub/Sub and Dataflow. The data source sometimes emits events that are delayed by several minutes due to network issues. Your pipeline must produce accurate aggregations (e.g., counts per minute) even for late data, but you also need to avoid waiting for a long time before emitting results. Which approach should you use?

A.Use processing-time windows and ignore the event timestamps entirely.
B.Use event-time processing with allowed lateness and a trigger that fires early to provide speculative results.
C.Use global windows and hold all data for 24 hours before processing to ensure completeness.
D.Use event-time processing and discard any data that arrives after the window ends.
AnswerB

Event-time processing aggregates by when events actually occurred, so delayed arrivals still land in the correct minute bucket. Allowed lateness retains those windows long enough to accept stragglers, while an early trigger emits speculative results promptly, satisfying the stem's dual constraint: accurate counts without long waits.

Why this answer

It uses event-time processing to handle late data via allowed lateness, combined with early triggers to emit speculative results before the window closes. This balances accuracy for delayed events with low latency for downstream consumers, which is a common requirement in streaming pipelines using Cloud Pub/Sub and Dataflow.

Exam trap

Google Cloud often tests the distinction between processing-time and event-time semantics, and the trap here is that candidates may choose processing-time windows (Option A) thinking they are simpler, not realizing they sacrifice correctness for late data.

How to eliminate wrong answers

Option A is wrong because processing-time windows ignore event timestamps entirely, so late-arriving data would be assigned to the wrong window, producing inaccurate aggregations. Option C is wrong because global windows with a 24-hour hold would cause unbounded latency and memory pressure, violating the requirement to avoid waiting a long time before emitting results. Option D is wrong because discarding late data after the window ends would lose delayed events, failing the requirement for accurate aggregations even with late data.

181
MCQmedium

The exhibit shows a Cloud Logging query result. A data engineer sees this log for a streaming Dataflow job. What is the most likely cause?

A.The job is experiencing network latency.
B.The job is using too much memory per worker.
C.The job has insufficient permissions to scale.
D.The job has reached the maximum number of workers allowed by the project quota.
AnswerD

Dataflow caps worker count at the project quota; once the autoscaler requests more workers than permitted, the job stalls and logs quota-related errors rather than processing backlog. This matches the observed log for a streaming pipeline that cannot scale further.

Why this answer

The log shows that the Dataflow job is not scaling up despite pending work. This typically occurs when the job has reached the maximum number of workers allowed by the project quota. Dataflow uses the Compute Engine default worker quota, and if the job attempts to exceed that limit, it will stop scaling and log messages indicating that it cannot add more workers.

Exam trap

Google Cloud often tests the distinction between resource quotas and permissions, so the trap here is that candidates confuse a quota limit (which is a hard resource cap) with an IAM permissions issue (which would produce a different error).

How to eliminate wrong answers

Option A is wrong because network latency would cause delays in data processing but would not prevent the job from scaling up; the log would show slow progress or timeouts, not a scaling block. Option B is wrong because excessive memory usage per worker would cause worker crashes or OOM errors, not a failure to scale; the job would still attempt to add workers. Option C is wrong because insufficient permissions to scale would result in an authorization error when trying to create new worker instances, not a quota-related log message; the error would reference IAM roles or service account permissions.

182
MCQmedium

After deploying a model to Vertex AI Endpoints, the prediction responses include unexpected data. The model returns logits instead of probabilities. What is the most likely cause?

A.The model was trained with different loss
B.The input data is scaled incorrectly
C.The endpoint is not properly configured
D.The model output is not post-processed
AnswerD

Raw logits are the model's unnormalised outputs; applying a softmax (or sigmoid for binary) transformation converts them to probabilities. The serving container returns the graph's final tensor unchanged, so without a post-processing step the endpoint emits logits, matching the unexpected-response symptom.

Why this answer

The most likely cause is that the model output is not post-processed. In Vertex AI Endpoints, models often output raw logits (unnormalized scores) from the final layer, and a softmax or sigmoid activation must be applied as a post-processing step to convert these logits into probabilities. Without this post-processing, the endpoint returns the raw logits, which is why the prediction responses contain unexpected data.

Exam trap

Google Cloud often tests the distinction between model training configurations and serving/post-processing steps, and the trap here is that candidates assume the endpoint or deployment configuration controls output formatting, when in fact the model's exported graph or serving function determines whether logits or probabilities are returned.

How to eliminate wrong answers

Option A is wrong because training with a different loss function (e.g., cross-entropy vs. mean squared error) does not directly cause the model to output logits instead of probabilities; the output layer's activation function (or lack thereof) determines whether outputs are logits or probabilities. Option B is wrong because incorrect input scaling would affect the prediction values (e.g., shifting or scaling them), but it would not change the fundamental nature of the output from logits to probabilities; the model would still output whatever its final layer produces. Option C is wrong because the endpoint configuration (e.g., machine type, traffic splitting, or model version) does not alter the model's output format; the endpoint simply serves the model's raw predictions as-is.

183
MCQhard

You are using BigQuery ML to train a matrix factorization model for a recommendation system. The training data consists of user-item interactions. You notice that the model is overfitting. Which of the following hyperparameter changes would most likely reduce overfitting?

A.Increase w_reg (regularization weight) from 0.1 to 0.5
B.Decrease w_reg (regularization weight) from 0.1 to 0.01
C.Increase num_factors from 10 to 20
D.Increase num_training_iterations from 10 to 20
AnswerA

w_reg controls L2 regularization strength on the learned factors. Raising it from 0.1 to 0.5 penalises large factor values more heavily, constraining model complexity and reducing overfitting on the sparse user-item interaction data described in the stem.

Why this answer

Increasing w_reg raises the L2 regularization penalty applied to the learned latent factor matrices, which shrinks factor magnitudes and constrains model complexity. In BigQuery ML's matrix factorization, w_reg directly controls how strongly the optimizer penalizes large weights, so moving from 0.1 to 0.5 tightens the fit and reduces overfitting on sparse user-item interaction data.

Exam trap

PDE often tests the misconception that adding more capacity (more factors or more iterations) improves a model, when in fact those changes worsen overfitting and only stronger regularization reduces it.

How to eliminate wrong answers

Option B is wrong because decreasing w_reg from 0.1 to 0.01 weakens regularization, allowing the model to fit training noise more aggressively and worsening overfitting. Option C is wrong because increasing num_factors from 10 to 20 expands the latent dimension, giving the model more capacity to memorize training interactions and typically increasing overfitting. Option D is wrong because increasing num_training_iterations from 10 to 20 only lets the optimizer run longer toward the same (or lower) training loss, which generally deepens overfitting rather than reducing it.

184
Multi-Selecthard

A company uses Cloud Build to deploy containerized applications. They want to ensure build and deployment quality. Which THREE steps should they include in their CI/CD pipeline? (Choose three.)

Select 3 answers
A.Scan container images for vulnerabilities using Container Analysis.
B.Run unit tests after deployment.
C.Deploy directly to production on every commit.
D.Use canary deployments with gradual traffic shifting.
E.Pin base image digests in Dockerfile.
AnswersA, D, E

Container Analysis scans built images for known vulnerabilities, gating deployment on findings and enforcing build quality before release. Integrating the scan into Cloud Build catches insecure base images and dependencies early, satisfying the pipeline's quality-assurance requirement.

Why this answer

Option A is correct because integrating Container Analysis into the Cloud Build pipeline scans container images for known vulnerabilities (CVEs) before they are pushed or deployed, which directly enforces build and deployment quality by catching insecure dependencies early. Option D is correct because canary deployments with gradual traffic shifting let you release a new revision to a small percentage of users first and monitor error rates and latency, so a bad build can be rolled back before it affects all production traffic. Option E is correct because pinning base image digests in the Dockerfile (e.g., FROM image@sha256:...) makes builds reproducible and immutable, preventing an unexpected or compromised 'latest' base image from silently changing the artifact.

Option B is wrong because unit tests must run before deployment, not after, to gate the pipeline on code correctness. Option C is wrong because deploying directly to production on every commit bypasses testing, review, and staged rollout, which is the opposite of ensuring quality.

Exam trap

Google often tests the misconception that unit tests can be run after deployment or that direct-to-production commits are acceptable in a quality-focused pipeline, when in fact both violate the principle of shifting left on quality and risk reduction.

185
MCQhard

A financial services company stores transaction logs in Cloud Storage. Regulatory requirements mandate that each object be retained for exactly seven years and that no user, including project owners, can delete or modify the objects during that period. The company wants to enforce this with minimal administrative overhead. What should they do?

A.Enable Uniform Bucket-Level Access and grant only the Storage Object Viewer role to all users.
B.Create a bucket with a retention policy of seven years and lock the policy.
C.Configure a Cloud Storage bucket lock using a bucket policy with a seven-year retention period.
D.Use Object Lifecycle Management to transition objects to Archive storage after 30 days and delete after seven years.
AnswerB

A locked retention policy enforces a minimum retention period during which objects cannot be deleted or overwritten, even by project owners. Locking the policy makes it permanent and irreversible, satisfying the WORM requirement. This is a native Cloud Storage feature that requires no custom code or external tooling, minimizing administrative overhead while meeting the exact seven-year retention mandate.

Why this answer

A locked retention policy on a Cloud Storage bucket enforces a minimum retention period and prevents deletion or modification of objects, even by project owners. Locking the policy makes it permanent, satisfying the seven-year WORM requirement. Lifecycle rules, IAM restrictions, or a nonexistent 'bucket policy' cannot provide the same immutable guarantee, making the locked retention policy the correct and minimal-overhead solution.

Exam trap

The trap here is confusing lifecycle management or IAM restrictions with true immutability, which only a locked retention policy provides.

186
MCQeasy

Your company wants to analyze real-time user clickstream data from a website. The data arrives as JSON messages via an HTTP endpoint. The pipeline should be able to handle spikes in traffic, provide low-latency insights, and store the raw data in a data lake for historical analysis. Which Google Cloud service should you use to ingest and process the streaming data?

A.Cloud Pub/Sub combined with Dataflow
B.Cloud Dataproc
C.Cloud Functions
D.Cloud IoT Core
AnswerA

Pub/Sub decouples ingestion from processing, absorbing traffic spikes via its globally distributed, auto-scaling message buffer, while Dataflow provides low-latency stream processing with exactly-once semantics and can write raw JSON to Cloud Storage for historical analysis.

Why this answer

Cloud Pub/Sub is the correct ingestion service because it provides a highly scalable, fully managed message queue that can handle traffic spikes by decoupling producers from consumers. Dataflow (Apache Beam) then processes the streaming data with low latency, supports exactly-once semantics, and can write raw data to a data lake like Cloud Storage for historical analysis. This combination meets all requirements: spike handling, low-latency insights, and raw data storage.

Exam trap

Google Cloud often tests the misconception that Cloud Functions can handle streaming ingestion due to its HTTP trigger, but its 9-minute timeout and lack of native streaming support make it unsuitable for high-throughput, low-latency pipelines.

How to eliminate wrong answers

Option B (Cloud Dataproc) is wrong because it is a managed Hadoop/Spark service designed for batch and stream processing but requires manual cluster management and autoscaling configuration, making it less suitable for handling unpredictable traffic spikes with low latency compared to the serverless Pub/Sub + Dataflow pipeline. Option C (Cloud Functions) is wrong because it is a lightweight, event-driven compute service with a maximum timeout of 9 minutes and limited throughput, making it unsuitable for high-volume, real-time streaming ingestion and processing. Option D (Cloud IoT Core) is wrong because it is specifically designed for ingesting data from IoT devices using MQTT/HTTP protocols, not for general web clickstream data from an HTTP endpoint, and it lacks the native streaming analytics capabilities needed for low-latency insights.

187
MCQmedium

A data engineer wants to create a data lake on Google Cloud for storing raw streaming data, then transform it into curated and processed zones for analytics. The data is in Avro format and will be queried by BigQuery. Which two services are MOST suitable as the primary storage and query interface?

A.Cloud Storage and BigQuery
B.Cloud Storage and Dataproc
C.Cloud Storage and Cloud SQL
D.Cloud Storage and Firestore
AnswerA

Cloud Storage provides the durable object store for raw Avro files, satisfying the data lake requirement, while BigQuery queries external data directly via its native Avro support and federated querying. This pairing separates cheap storage from serverless analytics, matching the raw-to-curated zone transformation the stem demands.

Why this answer

Cloud Storage is the most suitable primary storage for a data lake because it provides scalable, durable, and cost-effective object storage for raw Avro data. BigQuery is the ideal query interface because it can directly query Avro files stored in Cloud Storage using external tables, and it supports serverless analytics without needing to manage infrastructure.

Exam trap

Common misconception: candidates often think a processing engine like Dataproc is required to query Avro data in a data lake, but BigQuery can natively query Avro files stored in Cloud Storage without the need for intermediate processing.

How to eliminate wrong answers

Option B is wrong because Dataproc is a managed Spark/Hadoop service for batch processing, not a primary query interface for ad-hoc analytics; it would add unnecessary complexity and latency compared to BigQuery's direct Avro querying. Option C is wrong because Cloud SQL is a relational database for transactional workloads, not designed for large-scale analytics on Avro data in a data lake, and it cannot directly query Avro files. Option D is wrong because Firestore is a NoSQL document database for real-time applications, not suitable for analytical queries on large volumes of streaming data in Avro format.

188
MCQhard

A data science team deploys a TensorFlow image classification model to Vertex AI Prediction. The model performs well in offline evaluation but shows a 15% drop in accuracy in production. The production data distribution has shifted compared to the training data. The team needs to continuously monitor and retrain the model. Which solution is most appropriate for detecting drift and triggering retraining?

A.Enable Vertex AI Model Monitoring for feature drift; configure alerts to trigger a Vertex AI Pipelines retraining run.
B.Export production predictions to Cloud Logging, then use Log Analytics to compare distributions.
C.Store predictions in BigQuery and run scheduled SQL queries to detect drift; trigger retraining via Cloud Functions.
D.Use Cloud Monitoring to track prediction latency and error rates; manually retrain when errors increase.
AnswerA

Vertex AI Model Monitoring computes feature distribution skew and drift against a training baseline, raising alerts when thresholds are breached. Wiring those alerts to a Vertex AI Pipelines retraining run closes the loop automatically, directly addressing the production distribution shift and the need for continuous detection and retraining.

Why this answer

Vertex AI Model Monitoring is purpose-built for detecting feature drift in production ML models by comparing live inference data against a baseline distribution. When drift is detected, it can directly trigger a Vertex AI Pipelines retraining run, creating an automated, end-to-end MLOps loop that addresses the production accuracy drop without manual intervention.

Exam trap

Google Cloud often tests the distinction between operational monitoring (latency, errors) and data-quality monitoring (feature drift), leading candidates to mistakenly choose Cloud Monitoring (Option D) because they confuse production health metrics with model-specific distribution shifts.

How to eliminate wrong answers

Option B is wrong because exporting predictions to Cloud Logging and using Log Analytics for distribution comparison is a manual, ad-hoc approach that lacks native drift detection algorithms and automated retraining triggers, making it unsuitable for continuous monitoring. Option C is wrong because storing predictions in BigQuery and running scheduled SQL queries to detect drift requires custom statistical logic and does not leverage Vertex AI's built-in drift detection, alerting, or pipeline integration, leading to higher maintenance overhead. Option D is wrong because Cloud Monitoring tracks prediction latency and error rates, which are operational metrics, not feature distribution shifts; relying on error rates as a proxy for drift is indirect and unreliable, and manual retraining defeats the goal of continuous automation.

189
Drag & Dropmedium

Drag and drop the steps to create a Cloud Bigtable instance and table using the CLI into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Bigtable instances contain clusters; tables are created within instances and must have column families.

190
MCQmedium

A team wants to enforce data quality rules on BigQuery tables using Dataplex. They need to run column-level checks for null values and row-level checks for value ranges on a schedule. Which Dataplex feature should they use?

A.Dataplex Data Profiling
B.BigQuery stored procedures with scheduled queries
C.Dataplex Data Quality Tasks
D.Cloud DLP inspection jobs
AnswerC

Dataplex Data Quality Tasks run scheduled, rule-based checks directly against BigQuery tables, supporting both column-level null validation and row-level range conditions. This satisfies the stem's requirement for automated, recurring enforcement of data quality rules without external tooling, unlike profiling or discovery features that only observe metadata.

Why this answer

Dataplex Data Quality Tasks allow you to define and run data quality rules on BigQuery tables, including column-level checks (e.g., null checks) and row-level checks (e.g., value ranges). These tasks can be scheduled to run periodically, and they generate results that can be monitored. This is the native Dataplex feature for enforcing data quality.

Exam trap

PDE often tests the distinction between Dataplex Data Quality and Data Profiling, so candidates might choose Data Profiling for rule enforcement when it's actually for analysis.

How to eliminate wrong answers

Option A is wrong because Data Profiling analyzes data to discover statistics and patterns, but does not enforce rules or trigger alerts on violations. Option B is wrong because while stored procedures with scheduled queries can implement checks, they are not a Dataplex feature and require custom coding. Option D is wrong because Cloud DLP is for sensitive data discovery and de-identification, not for general data quality rules like null checks or value ranges.

191
MCQmedium

A company wants to move data from an on-premises MySQL database to BigQuery for analytics. They need to capture all changes (inserts, updates, deletes) in near real-time and also perform an initial historical load. Which approach meets these requirements with minimal operational overhead?

A.Use a Dataflow pipeline with a JDBC source to read the entire table periodically
B.Use Datastream to backfill historical data and then stream CDC changes to BigQuery
C.Use a one-time export to CSV and load into BigQuery, then set up a cron job to export incremental changes
D.Use Cloud SQL as an intermediary and enable binary logging, then stream to Pub/Sub via a custom connector
AnswerB

Datastream provides serverless change data capture from MySQL, streaming inserts, updates and deletes into BigQuery in near real-time, while its backfill capability performs the initial historical load. This satisfies both requirements without managing replication infrastructure, minimising operational overhead.

Why this answer

Datastream is correct because it provides serverless change data capture (CDC) from MySQL to BigQuery, supporting both historical backfill and continuous replication of inserts, updates, and deletes with minimal operational overhead. It reads the MySQL binary log (binlog) to stream changes in near real-time and can write directly to BigQuery or Cloud Storage.

Exam trap

PDE often tests the difference between batch and streaming ingestion; candidates incorrectly choose Dataflow or custom scripts for CDC, not realizing Datastream is the fully managed, serverless CDC service designed for minimal operational overhead.

How to eliminate wrong answers

Option A is wrong because a Dataflow JDBC pipeline reading entire tables periodically is batch-oriented, not near real-time, and does not capture deletes or updates efficiently; it also requires managing the pipeline and handling schema changes manually. Option C is wrong because one-time CSV export plus cron-based incremental exports is not real-time, cannot capture deletes, and requires custom scripting and scheduling, increasing operational overhead. Option D is wrong because using Cloud SQL as an intermediary and custom Pub/Sub connectors adds unnecessary complexity and operational burden; Datastream natively supports on-premises MySQL without needing Cloud SQL.

192
Multi-Selectmedium

A data engineer needs to monitor a Pub/Sub-based streaming pipeline. Which two Cloud Monitoring metrics should be used to detect a backlog of unprocessed messages? (Choose two.)

Select 2 answers
A.subscription/oldest_unacked_message_age
B.topic/byte_cost
C.subscription/num_undelivered_messages
D.topic/send_request_count
E.subscription/ack_message_count
AnswersA, C

oldest_unacked_message_age reports how long the oldest unacknowledged message has waited, directly exposing a growing backlog. Rising age means subscribers are not keeping pace with publishing, which is precisely the unprocessed-message condition the engineer must detect.

Why this answer

The 'subscription/num_undelivered_messages' metric shows the number of messages not yet acknowledged, and 'subscription/oldest_unacked_message_age' indicates how long messages have been waiting. Both help detect backlog.

193
Multi-Selecthard

You manage several Cloud Composer 2 environments that run production DAGs. You must define an alerting strategy that detects when a DAG run fails and when a task is stuck retrying for an unusually long time, using Cloud Monitoring. (Choose two.)

Select 2 answers
A.Create an uptime check against the Airflow web server URL and alert when the HTTP response is not 200.
B.Enable Cloud Audit Logs for the Composer API and alert when any environment update method is called.
C.Create a Monitoring alerting policy on the composer.googleapis.com/environment/healthy metric and notify when it drops below the threshold.
D.Create a Monitoring alerting policy on the airflow task instance duration or retry-related metric exposed through Cloud Monitoring for the environment.
E.Create a log-based alerting policy on the Composer airflow logs that matches the DAG run failure log entry and notifies the on-call channel.
AnswersD, E

Composer exposes Airflow metrics such as task instance duration and retry counts to Cloud Monitoring. An alerting policy on those metrics can detect a task whose runtime or retry behaviour exceeds a normal threshold, which is exactly the stuck-retrying condition described. It complements the failure alert by catching slow degradation before hard failure.

Why this answer

Composer surfaces Airflow execution data in two complementary ways: structured log entries in Cloud Logging and Airflow metrics in Cloud Monitoring. A log-based alerting policy catches the explicit DAG run failure event, while a metric-based policy on task duration or retries catches a task that is hanging or looping. Together they cover both the hard failure and the stuck-retrying condition, whereas health, uptime, and audit signals describe infrastructure and control-plane state rather than workload outcomes.

Exam trap

The trap here is reaching for environment health or web server uptime metrics, which measure whether Composer is running rather than whether your DAGs are succeeding.

194
Multi-Selecthard

A company is migrating their on-premises Apache Spark jobs to Google Cloud Dataproc. They want to minimize operational overhead and cost for jobs that run only a few times per day. Which TWO strategies should they adopt? (Choose TWO.)

Select 2 answers
A.Configure HDFS replication factor to 3 to ensure data durability during cluster restarts.
B.Rewrite the Spark jobs as Dataflow pipelines to take advantage of serverless processing.
C.Store all data in Cloud Storage instead of HDFS, and use the Cloud Storage connector to access it.
D.Create an ephemeral Dataproc cluster for each job and delete it after completion.
E.Use a small persistent cluster that runs continuously and submit jobs to it.
AnswersC, D

Correct. Storing data in Cloud Storage decouples storage from compute, enabling ephemeral clusters. The Cloud Storage connector provides Hadoop-compatible access, eliminating HDFS overhead and reducing cost because storage is billed separately and persists beyond cluster lifetime.

Why this answer

Option C is correct because storing data in Cloud Storage (GCS) and accessing it via the Cloud Storage connector decouples storage from compute, so ephemeral clusters can be deleted without losing data and you only pay for storage you actually use, which lowers cost and operational overhead. Option D is correct because ephemeral Dataproc clusters created per job and deleted on completion eliminate idle cluster costs and the patching/scaling overhead of managing long-running clusters, which fits jobs that run only a few times per day. Option A is not appropriate because HDFS replication factor 3 increases storage cost and is unnecessary when data should live in GCS for durability.

Option B is not appropriate because rewriting Spark jobs as Dataflow pipelines is a significant re-engineering effort and not required to reduce Dataproc operational overhead. Option E is not appropriate because a continuously running persistent cluster incurs idle compute costs and ongoing management overhead, contradicting the goal of minimizing cost and operations for infrequent jobs.

Exam trap

A common mistake is to think that a persistent cluster is needed for data durability or to avoid job startup latency. However, for jobs that run only a few times per day, ephemeral clusters with Cloud Storage are more cost-effective and operationally simpler.

195
MCQeasy

Your team wants to continuously monitor a deployed model's performance in production. They need to detect when the model's predictions become unreliable due to changes in the real world (e.g., new customer behavior). Which Vertex AI service should they use?

A.Vertex AI Explainable AI
B.Vertex AI Experiments
C.Vertex AI Model Monitoring
D.Vertex AI Prediction
AnswerC

Vertex AI Model Monitoring continuously evaluates deployed models against a training baseline, detecting drift and skew as real-world behaviour shifts. It satisfies the requirement to flag unreliable predictions caused by changing customer behaviour, without requiring manual retraining or custom metric pipelines.

Why this answer

Vertex AI Model Monitoring is the correct choice because it is specifically designed to continuously track a deployed model's prediction quality over time, detecting issues like data drift, feature drift, and prediction skew that indicate the model's reliability is degrading due to changes in the real world. It automatically compares incoming prediction data against a baseline training dataset and alerts when statistical distributions shift beyond configurable thresholds, enabling proactive retraining or intervention.

Exam trap

Google Cloud often tests the distinction between services that 'serve' predictions (Vertex AI Prediction) versus those that 'monitor' predictions (Vertex AI Model Monitoring), leading candidates to mistakenly choose the prediction service when the question asks about detecting unreliability.

How to eliminate wrong answers

Option A is wrong because Vertex AI Explainable AI provides feature attributions and explanations for individual predictions, but it does not continuously monitor model performance or detect drift in production. Option B is wrong because Vertex AI Experiments is used for tracking and comparing machine learning experiments during model development, not for monitoring deployed models in production. Option D is wrong because Vertex AI Prediction is the service that hosts and serves the model for online predictions, but it has no built-in monitoring capabilities for detecting performance degradation or data drift.

196
MCQmedium

A financial services company must analyze transaction data that includes customers' full names, account numbers, and home addresses. Regulations require that this personally identifiable information (PII) never be stored in raw form in their BigQuery analytics warehouse. The data engineering team plans to use Cloud Dataflow to read from a Pub/Sub topic and write to BigQuery. Which approach best satisfies the regulatory requirement while keeping the pipeline simple?

A.Apply a ParDo transform in Dataflow that uses the Cloud Data Loss Prevention (DLP) API to de-identify sensitive fields before writing to BigQuery.
B.Use BigQuery column-level security to restrict access to the PII columns, and grant permissions only to authorized analysts.
C.Enable BigQuery default encryption with a customer-managed encryption key (CMEK) so that the PII columns are encrypted at rest.
D.Load the raw data into a BigQuery table, then run a scheduled query that replaces PII columns with hashed values.
AnswerA

The Cloud DLP API is purpose-built for discovering and de-identifying sensitive data such as names, account numbers, and addresses. Calling it from a Dataflow ParDo transform lets you tokenize, mask, or encrypt PII in-flight, so only de-identified records reach BigQuery. This directly satisfies the regulation without adding storage layers or external systems.

Why this answer

De-identifying data in-flight with Cloud DLP inside a Dataflow pipeline prevents raw PII from ever reaching BigQuery storage. The DLP API can inspect and transform fields like names, account numbers, and addresses using techniques such as masking or tokenization. This meets the regulatory requirement at the earliest point, while keeping the pipeline serverless and managed.

Exam trap

The trap here is assuming that access controls or encryption at rest satisfy a requirement that PII must never be stored in raw form, when only de-identification before persistence actually removes the raw values.

197
MCQeasy

You are building a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. You need to ensure that no duplicates are written to BigQuery even if the pipeline retries messages. Which feature should you use?

A.Use a GroupByKey transform to deduplicate messages before writing to BigQuery.
B.Write to BigQuery using streaming inserts and then run a periodic deduplication job.
C.Use Pub/Sub's exactly-once delivery and rely on BigQuery's default streaming insert deduplication.
D.Enable exactly-once processing in Dataflow and use BigQuery's Storage Write API with exactly-once semantics.
AnswerD

Dataflow's exactly-once processing combined with the BigQuery Storage Write API's exactly-once delivery ensures that each message is written exactly once, even on retries. This is the most robust way to avoid duplicates in the target table.

Why this answer

The combination of Dataflow's exactly-once processing and the BigQuery Storage Write API's exactly-once semantics guarantees that each message is written once, even if the pipeline retries. This is the recommended approach for avoiding duplicates in streaming pipelines.

Exam trap

The trap here is relying on Pub/Sub's exactly-once delivery alone, which does not extend to the BigQuery sink and does not prevent duplicate writes.

198
MCQhard

In the Vertex AI Pipeline component YAML exhibit, the component is designed to evaluate a model and produce metrics. If the threshold_accuracy is set to 0.85, what is the expected behavior of this component?

A.It will output the evaluation metrics, and the pipeline can use them for conditional decisions
B.It will deploy the model if the accuracy meets the threshold
C.It will ignore the threshold_accuracy input if not provided
D.It will fail if the model accuracy is below 0.85
AnswerA

The component emits evaluation metrics as pipeline artefacts rather than enforcing the threshold itself; downstream conditional logic, such as a threshold comparison step, consumes those metrics to decide whether to promote the model. Setting threshold_accuracy to 0.85 therefore supplies the comparison value, not an automatic pass/fail gate.

Why this answer

In Vertex AI Pipelines, a component's YAML definition specifies inputs, outputs, and implementation. Setting `threshold_accuracy` to 0.85 defines a parameter that the component can use internally, but by itself it does not trigger deployment or cause failure. The component's expected behavior is to output evaluation metrics, and the pipeline can then use those metrics in conditional logic (e.g., via `Condition` or `if/else` tasks) to decide subsequent steps, such as model deployment or retraining.

Exam trap

Google Cloud often tests the misconception that setting a threshold in a component's YAML automatically enforces that threshold (e.g., causing failure or deployment), when in reality the YAML only defines the interface and the component's code must explicitly implement such logic.

How to eliminate wrong answers

Option B is wrong because Vertex AI Pipeline components do not inherently deploy models; deployment is a separate step typically handled by a deployment component or a pipeline condition that triggers a deployment task. Option C is wrong because if `threshold_accuracy` is not provided, the component will either use a default value defined in the YAML or fail validation, depending on whether the input is required; it does not simply ignore it. Option D is wrong because the component does not fail when accuracy is below the threshold; it merely outputs the metrics, and the pipeline logic (e.g., a conditional branch) must be explicitly configured to handle such cases.

199
MCQmedium

A company needs to process streaming sensor data from millions of devices with sub-second latency, apply transformations, and write results to BigQuery for real-time dashboards. The data volume varies, and they want to avoid managing servers. Which service should they use?

A.Cloud Data Fusion
B.Dataflow
C.Dataproc
D.Dataprep
AnswerB

Dataflow provides serverless, autoscaling stream processing with exactly-once semantics, satisfying the sub-second latency and variable-volume constraints. Its Apache Beam pipeline can transform sensor data and write directly into BigQuery via the built-in BigQueryIO connector, with no servers to manage.

Why this answer

Dataflow is Google Cloud's fully managed, serverless data processing service built on Apache Beam, designed for both batch and streaming pipelines with exactly-once processing semantics. It natively supports streaming ingestion from Pub/Sub or Kafka, applies transformations via Beam's windowing/triggers, and can write directly to BigQuery using the BigQueryIO connector with streaming inserts or Storage Write API. Its autoscaling (Horizontal Autoscaling and Streaming Engine) handles variable data volumes without server management, making it the only option that meets sub-second latency, serverless, and BigQuery streaming requirements simultaneously.

Exam trap

PDE often tests the distinction between serverless, streaming-native Dataflow and cluster-based or batch-oriented tools like Dataproc, Data Fusion, and Dataprep, causing candidates to pick Dataproc for 'streaming' because Spark Streaming sounds similar, or Data Fusion because it's an ETL tool.

How to eliminate wrong answers

Option A is wrong because Cloud Data Fusion is a GUI-based, code-free ETL orchestration tool built on CDAP; it is primarily batch-oriented, runs on a Dataproc cluster (so you still manage underlying infrastructure), and is not designed for sub-second streaming latency. Option C is wrong because Dataproc is a managed Hadoop/Spark cluster service where you still provision and size clusters (even with autoscaling, it's not serverless in the same sense) and Spark Streaming micro-batches typically have higher latency than Dataflow's per-record streaming. Option D is wrong because Dataprep by Trifacta is a data wrangling tool for interactive, visual preparation of batch data (often used with Dataflow under the hood for execution), not a streaming pipeline engine, and it lacks native sub-second streaming semantics.

200
Multi-Selectmedium

Which three of the following are valid BigQuery data loading methods? (Choose THREE.)

Select 3 answers
A.Data Transfer Service from Amazon S3
B.Using Cloud SQL to write to BigQuery
C.Direct file upload via Cloud Dataproc
D.Batch load from Cloud Storage
E.Streaming inserts using the legacy streaming API
AnswersA, D, E

BigQuery Data Transfer Service natively supports scheduled, recurring transfers from Amazon S3 into BigQuery, satisfying the stem's requirement for a valid loading method. It handles authentication, incremental refreshes and backfills automatically, unlike one-off manual loads, making it a legitimate ingestion path alongside batch loads and streaming inserts.

Why this answer

Option A is correct because BigQuery Data Transfer Service natively supports scheduled, managed transfers from Amazon S3 into BigQuery, making it a valid loading method. Option D is correct because batch loading from Cloud Storage is one of the primary and most common ways to load data into BigQuery, using load jobs that read from GCS URIs. Option E is correct because streaming inserts via the legacy streaming API (tabledata.insertAll) are a supported, documented method for loading data into BigQuery in real time.

Option B is not a valid loading method because Cloud SQL is a managed relational database service and does not write directly into BigQuery; data would need to be exported and then loaded via a supported path. Option C is not valid because Cloud Dataproc is a managed Spark/Hadoop service, and there is no direct file upload mechanism from Dataproc into BigQuery; Dataproc would typically write to Cloud Storage, which is then loaded into BigQuery.

Exam trap

Google often tests the distinction between data processing services (like Dataproc) and actual data loading methods, leading candidates to confuse a processing step with a direct ingestion path.

201
MCQhard

A company uses Eventarc to trigger a Cloud Run service when new objects appear in a GCS bucket. Recently, the Cloud Run service has been failing with 429 errors (too many requests) during high-velocity uploads. They need to handle the load without losing events. What should they do?

A.Use Cloud Functions instead of Cloud Run
B.Increase the maximum number of retries on the Eventarc trigger
C.Increase the Cloud Run service's request timeout
D.Configure the Eventarc trigger to send events to a Pub/Sub topic, and have the Cloud Run service pull from Pub/Sub
AnswerD

This decouples the event source from the consumer, allowing the Cloud Run service to process at its own pace and reducing 429 errors.

Why this answer

Sending events to a Pub/Sub topic decouples event production from consumption. Pub/Sub acts as a buffer that can absorb spikes in event volume, and the Cloud Run service can pull messages at its own pace, preventing 429 errors. This also ensures no events are lost, as Pub/Sub retains unacknowledged messages and retries delivery.

Exam trap

Google often tests the concept of decoupling event sources from consumers using a message queue or buffer like Pub/Sub, and the trap here is that candidates may think increasing retries or switching to Cloud Functions solves the load issue, when the real need is to absorb bursts via Pub/Sub.

How to eliminate wrong answers

Option A is wrong because Cloud Functions has similar concurrency limits and would also suffer from 429 errors under high load; it does not provide buffering. Option B is wrong because increasing retries on the Eventarc trigger only re-delivers failed events but does not address the root cause of the Cloud Run service being overwhelmed by too many concurrent requests. Option C is wrong because increasing the request timeout does not reduce the number of concurrent requests; it only allows longer processing time per request, which does not prevent the service from being overloaded.

202
MCQmedium

Refer to the exhibit. A developer sees this log entry when trying to get a prediction. What is the most likely cause?

A.The model ID is incorrect
B.The model version is not deployed
C.The endpoint does not exist
D.The project ID is wrong
AnswerB

The log indicates the prediction request reached the endpoint but no model version was available to serve it. Deploying the model version to the endpoint resolves this, since predictions require an active deployed version rather than merely a registered model.

Why this answer

The log entry indicates that the model version specified in the request is not currently deployed to the serving infrastructure. In Google Cloud's Vertex AI, a model version must be explicitly deployed to an endpoint before it can serve predictions; attempting to predict against a non-deployed version returns an error. This is the most likely cause because the error message directly references the model version's deployment status.

Exam trap

Google Cloud often tests the distinction between model registry operations (uploading, versioning) and serving operations (deploying, predicting), trapping candidates who assume any model version in the registry is automatically available for predictions.

How to eliminate wrong answers

Option A is wrong because an incorrect model ID would typically result in a 'Model not found' or 'Invalid model' error, not a deployment status error. Option C is wrong because a non-existent endpoint would produce a 'Endpoint not found' or 'Resource not found' error, not a version deployment issue. Option D is wrong because a wrong project ID would cause an authentication or permission error (e.g., 'Project not found' or 'Permission denied'), not a model version deployment error.

203
MCQeasy

Refer to the exhibit. A subscriber is unable to pull messages from the topic. What is the most likely cause?

A.The service account has the subscriber role but the topic is not configured correctly.
B.The service account needs roles/pubsub.viewer to list subscriptions.
C.No subscription has been created for the topic.
D.The service account lacks roles/pubsub.publisher.
AnswerC

Pub/Sub delivers messages only through subscriptions; a topic without any subscription retains messages but has no consumer to pull them. The subscriber's pull fails because no subscription exists to bind the topic to a receiving endpoint.

Why this answer

In Google Cloud Pub/Sub, a topic is a named resource to which messages are sent by publishers. Subscribers must create a subscription (pull or push) to receive messages from that topic. If no subscription exists, the subscriber cannot pull any messages because there is no delivery endpoint or pull queue attached to the topic.

Option C correctly identifies this missing subscription as the root cause.

Exam trap

Google Cloud often tests the distinction between topics and subscriptions, trapping candidates who assume that having a topic and a subscriber role is sufficient to receive messages, when in fact a subscription must be explicitly created.

How to eliminate wrong answers

Option A is wrong because the service account having the subscriber role (roles/pubsub.subscriber) is sufficient to pull messages; the topic configuration (e.g., schema, message retention) does not prevent pulling if a subscription exists. Option B is wrong because roles/pubsub.viewer only grants read access to list topics and subscriptions, not to pull messages; the subscriber role already includes the permission to list subscriptions (pubsub.subscriptions.list). Option D is wrong because roles/pubsub.publisher is required only to publish messages to a topic, not to pull messages; the subscriber role is the correct permission for pulling.

204
MCQeasy

A company is running a Cloud Dataflow streaming pipeline that aggregates events in 1-minute windows. They notice that the watermark is lagging significantly behind real-time. What is the most likely cause?

A.A hot key is causing data skew.
B.The window duration is too short.
C.The pipeline was recently updated.
D.The allowed lateness is set too high.
AnswerA

A hot key concentrates processing on one key, so a single worker stalls while others idle, delaying window closure and holding the watermark back. Data skew therefore explains the lagging watermark in the 1-minute windowed pipeline.

Why this answer

A hot key causes data skew, which means a disproportionate amount of data is assigned to a single key. In Cloud Dataflow, this leads to a single worker processing the bulk of the events, creating a processing bottleneck. The watermark, which tracks the progress of event-time processing, cannot advance until all data for a given window is processed, so the skewed key delays watermark progression significantly behind real-time.

Exam trap

Google Cloud often tests the misconception that watermark lag is caused by configuration settings like window duration or allowed lateness, rather than by data-level issues like hot keys that create processing bottlenecks.

How to eliminate wrong answers

Option B is wrong because a short window duration does not inherently cause watermark lag; it may increase computational overhead but does not prevent the watermark from advancing based on data arrival. Option C is wrong because a pipeline update (e.g., via a new job version) does not cause persistent watermark lag; it may cause a brief reprocessing delay but not a sustained lag. Option D is wrong because setting allowed lateness too high only affects how long the pipeline waits for late data after the watermark passes; it does not cause the watermark itself to lag behind real-time.

205
MCQeasy

You need to ingest data from a Cloud Storage bucket into BigQuery. The data is in Avro format and you want to minimize the time to insight. Which method should you use?

A.Use Cloud Data Fusion to read the Avro files and load them into BigQuery.
B.Use the BigQuery Streaming API to stream the Avro records from Cloud Storage.
C.Use a Dataflow pipeline to read the Avro files and write to BigQuery.
D.Use the BigQuery Load job with the Avro source format and enable autodetect schema.
AnswerD

Loading Avro files directly into BigQuery using a load job is the fastest and simplest method for batch ingestion. Avro is a supported format, and autodetect can infer the schema from the Avro file's embedded schema, eliminating manual schema definition. This approach minimizes time to insight because it avoids intermediate processing steps. It is ideal for one-time or scheduled batch loads from Cloud Storage.

Why this answer

The BigQuery Load job natively supports Avro and can automatically detect the schema from the Avro file's embedded schema. This makes it the fastest and simplest way to ingest Avro data from Cloud Storage, minimizing time to insight. Other methods like Dataflow, Streaming API, or Cloud Data Fusion add unnecessary complexity, latency, or cost for a straightforward batch load.

Exam trap

The trap here is overengineering the solution by choosing a pipeline or streaming service when a simple load job is sufficient and more efficient for batch Avro ingestion.

206
MCQhard

A data engineer wants to create a BigQuery external table that queries data stored in Parquet format in Cloud Storage without loading the data into BigQuery. Which approach is correct?

A.Use Cloud SQL federated query to read Parquet from GCS
B.Create an external table definition using a JSON schema file referencing Cloud Storage URIs
C.Use the bq load command with --source_format=PARQUET
D.Set up BigLake to create a BigQuery external table
AnswerB, D

External table definitions point to GCS, query data in place.

Why this answer

Creating an external table definition with a JSON schema file referencing Cloud Storage URIs is a valid direct way to query Parquet data without loading. BigLake (option D) is also a valid approach: BigLake tables are external tables over Cloud Storage that BigQuery can query without loading, with additional governance features. Therefore both B and D are correct.

Option A is Cloud SQL, not BigQuery; option C loads data into BigQuery.

Exam trap

The Google Professional Data Engineer exam often tests the distinction between loading data (bq load) and creating external tables (bq mk --external_table_definition), where candidates mistakenly choose the load command because they think it can also create external references.

How to eliminate wrong answers

Option A is wrong because Cloud SQL federated query is designed for querying Cloud SQL databases, not for reading Parquet files from Cloud Storage; it does not support Parquet as a data source. Option C is wrong because the `bq load` command loads data into BigQuery tables, not creating external tables; it physically moves the data into BigQuery storage, which contradicts the requirement to query without loading. Option D is wrong because BigLake is a separate service for managing data lakes with fine-grained access control, but creating a BigQuery external table directly via the console or API (with a schema definition) is the standard method; BigLake is not required for this simple external table use case.

207
MCQmedium

Your team runs a Cloud Composer 2 environment (composer-2.1.0-airflow-2.6.3) that executes a daily BigQuery ETL workflow. The workflow must not run on weekends. You want to implement this with minimal code and without modifying the DAG's task logic. What should you do?

A.Add a BranchPythonOperator that checks the day of the week and skips downstream tasks on weekends.
B.Set the DAG's schedule_interval to '@daily' and add a ShortCircuitOperator that stops execution on weekends.
C.Create a Cloud Scheduler job that triggers the Composer DAG only on weekdays via the Airflow REST API.
D.Set the DAG's schedule_interval to '0 2 * * 1-5' and set catchup=False.
AnswerD

The cron expression '0 2 * * 1-5' runs the DAG at 02:00 Monday through Friday, which excludes weekends. Setting catchup=False prevents backfilling missed runs. This directly satisfies the requirement to skip weekends without altering task logic, and it uses the native Airflow scheduling mechanism supported in Cloud Composer.

Why this answer

The correct approach is to use a cron expression that runs only Monday through Friday, such as '0 2 * * 1-5', combined with catchup=False to avoid backfilling. This leverages Airflow's native scheduling, requires no changes to task logic, and keeps the DAG simple. Other methods either add external dependencies or still trigger the DAG on weekends.

Exam trap

The trap here is assuming that a BranchPythonOperator or ShortCircuitOperator prevents the DAG from running on weekends, when in fact it only prevents downstream tasks from executing after the DAG has already been triggered.

208
MCQeasy

A media company stores video files in a Cloud Storage bucket and wants to serve them to users through a custom domain over HTTPS. The files are not sensitive, but the company wants to reduce egress cost and latency for users distributed globally. The company also wants to avoid managing SSL certificates manually. What should the data engineer recommend?

A.Create a Cloud CDN backend bucket with the Cloud Storage bucket as the origin and enable Google-managed SSL certificates on the load balancer.
B.Enable Object Versioning on the bucket and configure a lifecycle rule to move objects to Nearline after 30 days.
C.Set the bucket's default storage class to Standard and enable uniform bucket-level access.
D.Generate signed URLs for each video and distribute them to users through the application.
AnswerA

Cloud CDN caches content at edge locations, reducing latency and egress cost for globally distributed users. Using a backend bucket with the Cloud Storage bucket as origin, combined with a global external Application Load Balancer and Google-managed certificates, provides HTTPS on a custom domain without manual certificate management.

Why this answer

Cloud CDN with a backend bucket caches video content at Google's edge locations, lowering latency and egress for global users. Pairing it with a global external Application Load Balancer and Google-managed SSL certificates delivers HTTPS on a custom domain without the operational burden of managing certificates.

Exam trap

The trap here is focusing on storage class or access control features, which affect cost and permissions but do nothing for global delivery latency or custom domain HTTPS.

209
MCQmedium

A financial services company uses Cloud Composer to orchestrate daily batch jobs. One job extracts data from MongoDB to Cloud Storage, then loads into BigQuery, and finally runs a Dataflow pipeline for aggregations. The Dataflow job fails intermittently. They want to automatically restart only the failed Dataflow job without re-running the earlier extraction and load. Which Airflow operator configuration should they use?

A.Implement a SlaMiss sensor
B.Use a DAG with depends_on_past=True
C.Set retries=2 on the Dataflow operator
D.Set trigger_rule='one_success' for downstream tasks
AnswerC

Airflow retries apply at task level, so setting retries=2 on the Dataflow operator restarts only that failed task while upstream extract and load tasks remain successful and are not re-run. This satisfies the requirement to avoid repeating earlier extraction and loading.

Why this answer

Setting retries=2 on the Dataflow operator instructs Airflow to automatically restart only that specific task upon failure, without affecting upstream tasks (MongoDB extraction, BigQuery load). This isolates the retry to the Dataflow job, preserving the earlier completed work and avoiding redundant data movement.

Exam trap

Google Cloud often tests the distinction between task-level retry mechanisms and dependency/trigger rules, so the trap here is confusing `retries` (which restarts the failed task) with `trigger_rule` or `depends_on_past` (which only affect task scheduling or downstream execution).

How to eliminate wrong answers

Option A is wrong because SlaMiss sensors are used to detect when tasks have not completed within a defined SLA window, not to trigger automatic retries of failed tasks. Option B is wrong because depends_on_past=True enforces sequential execution order across DAG runs (e.g., today’s task waits for yesterday’s success), but does not provide automatic retry on failure within the same run. Option D is wrong because trigger_rule='one_success' controls downstream task execution based on upstream task outcomes (e.g., if one upstream succeeds, proceed), but does not restart a failed task; it only affects task dependencies.

210
MCQeasy

A data engineer notices that Spark jobs on the Dataproc cluster shown often fail with executor lost errors. What is the most likely reason?

A.All 10 workers are preemptible and can be reclaimed by Compute Engine at any time.
B.The master node has only 4 vCPUs, which may be insufficient for job coordination.
C.The cluster is in a single zone, so a zone failure could cause all workers to shut down.
D.Autoscaling is enabled and scaling down is causing workers to be removed during job execution.
AnswerA

Preemptible VMs are reclaimed by Compute Engine within 24 hours, and typically much sooner under capacity pressure. With all 10 workers preemptible, any reclamation kills executors mid-task, producing the executor lost failures described. The stem's constraint — every worker being preemptible — removes any stable capacity to absorb those losses.

Why this answer

Preemptible VMs in Google Compute Engine can be terminated at any time due to resource contention or other factors, with only 30 seconds notice. If all 10 worker nodes are preemptible, Spark executors running on them will be frequently lost, causing job failures. This is the most direct cause of 'executor lost' errors in a Dataproc cluster.

Exam trap

The trap here is that candidates may overlook the 'all 10 workers are preemptible' detail and instead focus on common misconfigurations like single-zone risk or autoscaling, but the explicit mention of preemptible VMs is the key indicator of frequent, unpredictable executor loss.

How to eliminate wrong answers

Option B is wrong because the master node's vCPUs (4) are typically sufficient for job coordination; executor lost errors are not caused by insufficient master resources but by worker instability. Option C is wrong because a single-zone cluster does not cause frequent executor losses; zone failures are rare and would cause complete cluster failure, not intermittent executor lost errors. Option D is wrong because autoscaling removes workers gracefully, allowing Spark to reschedule tasks before termination; it does not cause the abrupt 'executor lost' errors seen here.

211
Multi-Selectmedium

A data engineer is migrating on-premises Hadoop jobs to Dataproc. Which TWO considerations are important?

Select 2 answers
A.Use Preemptible VMs for worker nodes to reduce cost
B.Use Cloud Storage instead of HDFS for data storage
C.Avoid using Cloud Storage connector to prevent overhead
D.Keep HDFS for better performance
E.Use on-demand VMs for master node to ensure availability
AnswersA, B

Preemptible VMs cost far less than standard instances, and Dataproc tolerates their eviction because HDFS is replaced by Cloud Storage. This satisfies the cost-reduction constraint while keeping jobs resilient, provided workers are not running single-node or stateful workloads.

Why this answer

Option A is correct because Dataproc supports preemptible VMs (now called Spot VMs) for worker nodes, which can significantly reduce compute costs for fault-tolerant Hadoop workloads, though they may be reclaimed; Dataproc handles this by replacing preempted workers. Option B is correct because Cloud Storage (GCS) is the recommended storage layer for Dataproc, decoupling storage from compute, allowing clusters to be shut down without data loss, and providing higher durability and scalability than HDFS. Option C is incorrect because the Cloud Storage connector is essential for reading and writing data between Dataproc and GCS, and it does not introduce prohibitive overhead; avoiding it would prevent access to GCS.

Option D is incorrect because keeping HDFS ties data to the cluster lifecycle, defeating the benefits of cloud elasticity and durability that GCS provides. Option E is incorrect because while master node availability matters, using on-demand VMs for the master is not a specific migration consideration highlighted here, and Dataproc already uses standard VMs for masters by default; the key cost and storage considerations are preemptible workers and GCS.

Exam trap

A common misconception is that HDFS must be retained for performance in the cloud, but the correct approach is to use Cloud Storage for data storage and Preemptible VMs for cost savings, while the Cloud Storage connector is a required component, not an overhead to avoid.

212
MCQeasy

You need to load a 2 TB CSV file from Cloud Storage into BigQuery. The CSV has a header row and uses a comma delimiter. You want to minimize cost and ensure the load completes quickly. Which method should you use?

A.Run a `bq load` command with `--source_format=CSV` and `--skip_leading_rows=1`.
B.Stream the data using the BigQuery Storage Write API.
C.Use the BigQuery web UI to upload the file directly.
D.Use the BigQuery Data Transfer Service to schedule a recurring load.
AnswerA

The `bq load` command can load large files from Cloud Storage into BigQuery. Specifying `--source_format=CSV` and `--skip_leading_rows=1` handles the header row. This method is cost-effective (no data egress) and leverages BigQuery's parallel load capabilities for fast ingestion.

Why this answer

For large files in Cloud Storage, the most efficient and cost-effective method is a batch load using the `bq load` command or a load job. It supports CSV format, handles header rows with `--skip_leading_rows`, and parallelizes ingestion. This avoids the limitations of UI uploads and the overhead of streaming or scheduled transfers.

Exam trap

The trap here is choosing streaming or the web UI for a large file; batch loading from Cloud Storage is the recommended approach for bulk data.

213
MCQeasy

A data engineer needs to process large CSV files (hundreds of GB) stored in Cloud Storage using Spark on a Dataproc cluster. The job performs a series of transformations and aggregations. Which configuration is most cost-effective and operationally efficient?

A.Use a cluster with 10 high-memory (n1-highmem-8) VMs as workers to improve shuffle performance.
B.Use a cluster with a standard master node and 10 preemptible worker nodes (n1-standard-4).
C.Use a single-node cluster with a high-memory machine type.
D.Use a cluster with 10 standard (n1-standard-4) VMs as master and worker nodes, all non-preemptible.
AnswerB

Preemptible workers cost far less than standard VMs, and ten n1-standard-4 nodes provide enough parallelism for hundreds of GB of CSV transformations. A standard master preserves cluster stability, making this the cost-effective, operationally efficient configuration.

Why this answer

Preemptible workers are significantly cheaper (about 80% discount) and ideal for batch processing of large CSV files where fault tolerance is built into Spark via RDD lineage. Using standard nodes for the master ensures cluster stability, while preemptible workers handle the distributed transformations and aggregations cost-effectively. This configuration balances cost and operational efficiency for ephemeral, fault-tolerant workloads.

Exam trap

Google Cloud often tests the misconception that preemptible VMs are unreliable for all workloads, but in Spark batch processing with fault tolerance, they are both cost-effective and operationally efficient, unlike stateful or latency-sensitive applications.

How to eliminate wrong answers

Option A is wrong because using high-memory VMs (n1-highmem-8) for all workers increases cost unnecessarily; shuffle performance is better addressed by tuning Spark parameters (e.g., spark.shuffle.partitions) and using SSDs, not by over-provisioning memory. Option C is wrong because a single-node cluster cannot process hundreds of GB of data efficiently due to lack of parallelism and memory constraints, and it violates the distributed processing paradigm of Spark. Option D is wrong because using all non-preemptible standard VMs (n1-standard-4) for both master and workers eliminates the cost savings of preemptible instances, and having a separate master node is unnecessary for small clusters—the driver can run on a worker—but the main issue is the higher cost without fault-tolerance benefits.

214
MCQhard

You are designing a Dataflow pipeline that needs to exactly-once process events from Pub/Sub and write to BigQuery using the Storage Write API. The pipeline may restart and could reprocess some messages. What setting ensures exactly-once semantics for the output?

A.Use the legacy streaming inserts with insertId for deduplication
B.Use at-least-once delivery on Pub/Sub and idempotent writes to BigQuery
C.Use the Storage Write API in buffered mode with deduplication logic
D.Use the Storage Write API in committed mode and enable exactly-once semantic in Dataflow
AnswerD

Committed mode on the Storage Write API makes each write atomic and idempotent via stream offsets, and enabling exactly-once in Dataflow deduplicates retried bundles. Together they prevent duplicate rows when the pipeline restarts and reprocesses Pub/Sub messages.

Why this answer

The BigQuery Storage Write API in committed mode, combined with Dataflow's exactly-once processing, provides exactly-once semantics for writes to BigQuery. Committed mode uses a stream-based protocol where records are written and committed atomically, and Dataflow's exactly-once mode ensures that on pipeline restart, records are not duplicated. This is the recommended configuration for exactly-once streaming into BigQuery.

Exam trap

The trap is choosing legacy streaming inserts with insertId, which only provides best-effort deduplication, instead of the Storage Write API committed mode with Dataflow exactly-once, which is the actual mechanism for exactly-once semantics.

How to eliminate wrong answers

Option A is wrong because legacy streaming inserts with insertId provide best-effort deduplication, not true exactly-once semantics, and are deprecated for new pipelines. Option B is wrong because at-least-once delivery plus idempotent writes does not guarantee exactly-once output; duplicates can still occur if idempotency is imperfect. Option C is wrong because buffered mode does not provide the same exactly-once guarantees as committed mode and requires manual deduplication logic, which is not a built-in exactly-once setting.

215
MCQmedium

Your team runs a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. During a deployment, the pipeline is stopped and later restarted with the same job name but without draining. After the restart, downstream reports show duplicate rows for events that were processed just before the stop. You need the pipeline to resume without reprocessing already-published messages. What should you do?

A.Increase the number of workers with --maxNumWorkers before restarting the pipeline.
B.Drain the pipeline before stopping it, then start the replacement pipeline from the saved state.
C.Enable --enableStreamingEngine on the restarted pipeline to preserve message state.
D.Restart the pipeline using the --update option so the transform graph is replaced in place.
AnswerB

Draining a streaming Dataflow job stops ingestion of new data, finishes processing the in-flight elements, and commits the Pub/Sub acknowledgements and BigQuery writes. When the replacement job starts, it resumes from the committed state rather than replaying messages that were already acknowledged, which prevents the duplicate rows observed in the reports.

Why this answer

Stopping a streaming job without draining leaves in-flight Pub/Sub messages unacknowledged, so they are redelivered when a replacement job starts, producing duplicates. Draining first lets the pipeline finish processing and commit those acknowledgements and BigQuery writes, so the replacement resumes from a consistent state and avoids reprocessing the same events.

Exam trap

The trap here is assuming that restarting with the same job name or adding the update flag preserves exactly-once progress, when only draining commits in-flight Pub/Sub acknowledgements.

216
MCQmedium

A company has a Cloud SQL for PostgreSQL instance that must be highly available across zones with automatic failover. They also need a read replica for reporting workloads. Which configuration should they use?

A.Enable high availability (regional) and create a read replica
B.Deploy Cloud SQL in multi-region mode
C.Enable automatic backups and point-in-time recovery
D.Create a cross-region replica and use it for failover
AnswerA

A regional configuration replicates the primary across zones and provides automatic failover, while a separate read replica offloads reporting queries without competing with transactional traffic. Both stated requirements — cross-zone high availability and reporting capacity — are satisfied simultaneously.

Why this answer

Cloud SQL for PostgreSQL supports regional high availability (HA) with synchronous replication across two zones, ensuring automatic failover without data loss. Additionally, you can create a read replica in a different zone or region to offload reporting workloads without impacting the primary instance's performance. This combination meets both the high availability and read replica requirements.

Exam trap

The trap here is confusing high availability (automatic zonal failover) with disaster recovery (manual cross-region promotion), and assuming that a read replica can serve as a failover target without understanding that Cloud SQL read replicas use asynchronous replication and are not suitable for automatic failover.

How to eliminate wrong answers

Option B is wrong because Cloud SQL does not support a 'multi-region mode'; it offers regional HA (zonal failover) and cross-region replicas, but not a multi-region deployment like Spanner. Option C is wrong because automatic backups and point-in-time recovery provide data protection and restore capabilities, but they do not provide automatic failover or a read replica for reporting. Option D is wrong because a cross-region replica can be promoted for failover, but it is not designed for automatic failover; promoting a cross-region replica is a manual process and does not provide the same synchronous replication and automatic failover as regional HA.

217
Multi-Selectmedium

Which TWO statements are true about BigQuery Data Transfer Service? (Choose 2)

Select 2 answers
A.It supports data transformation during transfer.
B.It can transfer data from Cloud SQL to BigQuery directly.
C.It is only available in the US and EU regions.
D.It supports scheduled transfers from Amazon S3 and Redshift.
E.It can transfer data from Google Ads, YouTube, and Google Ad Manager into BigQuery.
AnswersD, E

BigQuery Data Transfer Service provides connectors that pull data from Amazon S3 and Amazon Redshift on scheduled, managed transfers into BigQuery. This cross-cloud ingestion capability satisfies the requirement, distinguishing it from services limited to Google-native sources.

Why this answer

Option D is correct because BigQuery Data Transfer Service provides built-in connectors for Amazon S3 and Amazon Redshift, enabling recurring, scheduled (and on-demand) transfers of external data into BigQuery without custom code. Option E is correct because the service also offers first-party Google connectors, including Google Ads, YouTube (Channel/Content Owner reports), and Google Ad Manager, which load those marketing datasets into BigQuery on a schedule. Option A is not correct because the Data Transfer Service is designed for ingestion/movement, not in-flight transformation; transformations are typically done after loading using BigQuery SQL, views, or scheduled queries.

Option B is not correct because Cloud SQL is not a native Data Transfer Service source; moving Cloud SQL data into BigQuery generally requires tools such as Datastream, the Cloud SQL federated query/EXTERNAL_QUERY, or export/load pipelines. Option C is not correct because BigQuery Data Transfer Service is available in many more locations than just US and EU, and its regional availability follows BigQuery dataset locations rather than being limited to those two regions.

Exam trap

A common misconception is that BigQuery Data Transfer Service can perform ETL transformations during transfer, but it is strictly an EL (Extract and Load) service without built-in transformation capabilities.

218
MCQmedium

A company is using Pub/Sub to ingest clickstream events. They need to ensure that events are delivered to a subscriber at least once, but duplicates can be tolerated. They also need to filter events by type before processing. Which subscription configuration should be used?

A.Pull subscription with exactly-once delivery enabled
B.Push subscription with no filter
C.Pull subscription with a filter on event type attribute
D.Push subscription with a dead letter topic
AnswerC

A pull subscription supports attribute filters, so Pub/Sub discards non-matching messages server-side before delivery, satisfying the filtering requirement. Pull delivery retains at-least-once semantics, meaning occasional duplicates may reach the subscriber, which the stem explicitly tolerates.

Why this answer

A pull subscription with a filter on the event type attribute satisfies both requirements: Pub/Sub's at-least-once delivery is the default for pull subscriptions, so duplicates are tolerated, and message filtering via subscription filter expressions (on attributes like event type) prevents unwanted events from being delivered. This is the only option that combines at-least-once semantics with server-side filtering.

Exam trap

PDE often tests the confusion between at-least-once and exactly-once delivery semantics, tricking candidates into selecting exactly-once when the scenario explicitly says duplicates are acceptable.

How to eliminate wrong answers

Option A is wrong because exactly-once delivery is the opposite of at-least-once — it eliminates duplicates rather than tolerating them, and it is not the requirement here. Option B is wrong because a push subscription with no filter provides no filtering capability, so all events would be delivered regardless of type. Option D is wrong because a dead letter topic handles undeliverable messages, not event-type filtering, and push subscriptions do not address the filtering requirement.

219
MCQmedium

A logistics company ingests 5 TB of shipment telemetry per day into a BigQuery table named events_raw. Analysts run exploratory queries that scan the full table, which is now 900 TB, and costs are rising quickly. Most queries filter on the event_date column, which is a DATE field, and on the device_id column, but the table is not partitioned or clustered. The company wants to reduce bytes scanned without changing the query patterns or losing any historical events. Which combination of BigQuery table settings should the data engineer implement?

A.Partition events_raw by event_date and cluster by device_id.
B.Set a table expiration of 30 days on events_raw and rely on time travel for older data.
C.Cluster events_raw by event_date and device_id without partitioning.
D.Create a materialized view that aggregates events_raw by event_date and device_id.
AnswerA

Partitioning on event_date lets the query engine prune partitions when a date filter is present, while clustering on device_id sorts and colocation-stores rows within each partition so filters on device_id scan fewer blocks. Together they reduce bytes billed for the filters the analysts already use, and no rows are removed, so all historical events remain available.

Why this answer

Partitioning by a DATE column enables partition pruning whenever queries filter on that column, and clustering on a high-cardinality filter column such as device_id reduces the blocks scanned inside each partition. Because the requirement is to keep all events and not rewrite queries, this physical layout change is the right fit. Aggregation views and expiration policies do not address the full-scan cost while preserving raw history.

Exam trap

The trap here is assuming clustering alone gives the same pruning as partitioning, when only partitioning can eliminate entire date ranges from a scan.

220
MCQmedium

A machine learning team wants to deploy a new model version for canary testing, where only 5% of traffic is routed to the new version. Which Vertex AI endpoint configuration supports this?

A.Have the client application randomly select which model to call with 5% probability.
B.Deploy the new version to a separate endpoint and direct 5% of users via a load balancer.
C.Configure the endpoint with traffic split: 95% to old version, 5% to new version.
D.Use an A/B testing framework outside of Vertex AI to compare results.
AnswerC

Vertex AI endpoints support traffic splitting by assigning percentage weights to deployed model versions. Setting 95% to the existing version and 5% to the new one routes exactly the required canary share, enabling gradual validation before full rollout.

Why this answer

Vertex AI endpoints natively support traffic splitting, allowing you to route a specified percentage of requests to different model versions deployed on the same endpoint. By configuring a traffic split of 95% to the old version and 5% to the new version, you can perform canary testing without additional infrastructure or client-side logic. This is the correct and simplest approach within Vertex AI.

Exam trap

The trap here is that candidates may think canary testing requires external tools or client-side logic, but Vertex AI's built-in traffic splitting is the intended and simplest method for this purpose.

How to eliminate wrong answers

Option A is wrong because it requires modifying the client application to implement random selection, which is error-prone, not managed by Vertex AI, and does not provide centralized traffic management or monitoring. Option B is wrong because deploying to a separate endpoint and using an external load balancer adds unnecessary complexity and bypasses Vertex AI's built-in traffic splitting capabilities, which are designed for this exact use case. Option D is wrong because using an external A/B testing framework outside Vertex AI would require custom integration and does not leverage Vertex AI's native traffic split feature, which is simpler and more reliable for canary deployments.

221
MCQeasy

A company has a BigQuery dataset containing sensitive customer data. They want to share a subset of this data with external partners, ensuring that partners can only see specific columns and rows. Which BigQuery feature should they use?

A.Materialized views
B.Authorized views
C.Clustered tables
D.Dataset-level access controls
AnswerB

Authorized views grant external partners query access to a view while restricting them to the columns and rows the view exposes, without granting underlying table access. This directly satisfies the requirement to share only a subset of sensitive customer data.

Why this answer

Authorized views in BigQuery allow a view to be created in one dataset that reads from a source dataset, and then access to the view is granted to external partners without granting access to the underlying source data. This enables column- and row-level filtering (via the view's SQL) so partners see only the permitted subset. This is the standard BigQuery pattern for sharing restricted data.

Exam trap

PDE often tests the misconception that dataset-level permissions or materialized views provide fine-grained sharing — candidates pick dataset access controls when the requirement is column/row filtering, which only authorized views deliver.

How to eliminate wrong answers

Option A is wrong because materialized views are for performance optimization (precomputed results) and do not by themselves provide a secure sharing boundary; they still require access to underlying data. Option C is wrong because clustered tables improve query performance by physically sorting data, not access control. Option D is wrong because dataset-level access controls grant access to the entire dataset, not a filtered subset, so partners would see all columns and rows.

222
Multi-Selecteasy

A data engineer is using Vertex AI Workbench to develop a custom ML model. They want to store and version datasets, track experiments, and register models. Which three Vertex AI services should they use? (Choose THREE)

Select 3 answers
A.Vertex AI Model Registry
B.Vertex AI Dataset
C.Vertex AI Feature Store
D.Vertex AI Matching Engine
E.Vertex AI Experiments
AnswersA, B, E

Vertex AI Model Registry provides a centralised repository for managing model versions, storing artefacts, and tracking lineage. It satisfies the requirement to register models by assigning each trained model a versioned entry with metadata, enabling deployment and monitoring. Combined with Datasets and Experiments, it completes the MLOps workflow the engineer needs.

Why this answer

Vertex AI Model Registry (A) is correct because it provides a central repository to manage the lifecycle of ML models, including versioning, tracking lineage, and deploying models to endpoints, which directly satisfies the requirement to register models. Vertex AI Dataset (B) is correct because it lets the engineer create managed, versioned datasets from tabular, image, text, or video data, fulfilling the need to store and version datasets. Vertex AI Experiments (E) is correct because it records and compares experiment runs, including parameters, metrics, and artifacts, which satisfies the requirement to track experiments.

Vertex AI Feature Store (C) is not appropriate here because it is designed for serving and sharing online/offline feature values, not for dataset versioning, experiment tracking, or model registration. Vertex AI Matching Engine (D) is also not appropriate because it is a vector similarity search service for embeddings, unrelated to the requested dataset, experiment, and model management tasks.

Exam trap

PDE often tests the specific purpose of each Vertex AI service; candidates might confuse Feature Store with Dataset or Matching Engine with Model Registry.

223
Multi-Selectmedium

A company deploys a TensorFlow model on Vertex AI for online predictions. They want to monitor model performance in production to detect degradation. Which TWO practices should they implement? (Choose 2.)

Select 2 answers
A.Use a separate endpoint for shadow testing new model versions.
B.Log prediction requests and responses to Cloud Logging and analyze distribution metrics.
C.Set up Cloud Monitoring alerts for high prediction latency.
D.Schedule daily retraining of the model regardless of monitoring alerts.
E.Enable Vertex AI Model Monitoring for feature drift and skew detection on the deployed model.
AnswersB, E

Analyzing request distributions can detect changes in input data patterns that may affect model performance.

Why this answer

Logging prediction requests and responses to Cloud Logging allows you to analyze distribution metrics (e.g., mean, variance, quantiles) over time. This enables detection of data drift or performance degradation by comparing live distributions against baseline distributions, which is a standard monitoring practice for production ML models.

Exam trap

Google Cloud often tests the distinction between monitoring for model degradation (data drift/skew) versus monitoring for operational issues (latency, errors), leading candidates to confuse infrastructure alerts with model performance monitoring.

224
MCQeasy

A team needs to store transactional data for an e-commerce application that requires ACID transactions, automatic backups, and point-in-time recovery. The expected workload is under 10,000 QPS. Which database should they choose?

A.Cloud Bigtable
B.Cloud SQL
C.Firestore
D.Cloud Spanner
AnswerB

Cloud SQL delivers the required ACID transactions, automatic backups and point-in-time recovery, and comfortably handles workloads below 10,000 QPS. Its managed MySQL, PostgreSQL and SQL Server engines suit transactional e-commerce data, unlike analytics-oriented or NoSQL alternatives that lack full relational ACID guarantees.

Why this answer

Cloud SQL is the correct choice because it provides fully managed relational databases (MySQL, PostgreSQL, SQL Server) with built-in ACID transaction support, automated backups, and point-in-time recovery (PITR) via binary logs or write-ahead logs. The workload of under 10,000 QPS is well within Cloud SQL's performance envelope, making it a cost-effective and operationally simple solution for transactional e-commerce data.

Exam trap

The trap here is that candidates often choose Cloud Spanner for any workload requiring ACID transactions and high availability, overlooking that Cloud SQL is sufficient and more cost-effective for sub-10,000 QPS workloads, and that Cloud Spanner's global distribution and strong consistency come with a significant price premium.

How to eliminate wrong answers

Option A is wrong because Cloud Bigtable is a NoSQL, wide-column database designed for high-throughput analytical workloads (millions of QPS) and does not support ACID transactions or SQL queries, making it unsuitable for transactional e-commerce data. Option C is wrong because Firestore is a NoSQL document database that, while supporting transactions, does not offer the full ACID guarantees across multiple documents in the same way as a relational database, and its automatic backup and PITR capabilities are limited compared to Cloud SQL. Option D is wrong because Cloud Spanner is a globally distributed, horizontally scalable relational database that supports ACID transactions and PITR, but it is overkill and significantly more expensive for a workload under 10,000 QPS, which can be handled more cost-effectively by Cloud SQL.

225
MCQmedium

After deploying a model, the team notices that predictions are significantly different from training data distribution. What should they do?

A.Update the model endpoint
B.Review the training data pipeline
C.Set up Vertex AI Model Monitoring for skew detection
D.Retrain the model with new data
AnswerC

Training-serving skew arises when live feature distributions diverge from those seen during training. Vertex AI Model Monitoring computes skew metrics by comparing production traffic against the training baseline, directly satisfying the requirement to detect distributional divergence and alert the team before predictions degrade further.

Why this answer

Vertex AI Model Monitoring is specifically designed to detect skew between training data and serving data, including prediction drift. When predictions differ significantly from the training distribution, this indicates a skew or drift issue that Model Monitoring can alert on, enabling proactive investigation. Updating the endpoint or retraining without diagnosis would not address the root cause, and reviewing the pipeline alone does not provide ongoing detection.

Exam trap

Google Cloud often tests the distinction between reactive troubleshooting (reviewing pipelines, retraining) and proactive monitoring (skew detection), tempting candidates to choose a fix like retraining instead of the monitoring solution that detects the issue first.

How to eliminate wrong answers

Option A is wrong because updating the model endpoint does not diagnose or resolve the distribution mismatch; it only changes the serving target without addressing the underlying data or model behavior. Option B is wrong because reviewing the training data pipeline is a reactive, one-time investigation step, whereas the question describes a deployed model scenario where continuous monitoring is needed to detect and alert on skew in real time. Option D is wrong because retraining with new data without first understanding the cause of the skew may introduce new biases or fail to fix the issue; monitoring should be used to detect and diagnose before retraining.

Page 2

Page 3 of 10

Page 4

All pages