Courseiva

Google Professional Data Engineer (PDE) — Questions 376–450

747 questions total · 10pages · All types, answers revealed

Page 5

Page 6 of 10

Page 7
376
Multi-Selectmedium

You are designing a streaming Dataflow pipeline that processes high-throughput data. Which two features can help minimize cost? (Choose TWO.)

Select 2 answers
A.Enable autoscaling based on CPU utilization
B.Use batch loads to BigQuery for streaming inserts
C.Enable Streaming Engine to decouple compute and storage
D.Use preemptible VMs for all workers
E.Use a global window and batch output to BigQuery every hour
AnswersA, C

Autoscaling adjusts the number of workers to meet demand, avoiding over-provisioning and reducing cost.

Why this answer

Enabling autoscaling based on CPU utilization allows the Dataflow pipeline to dynamically adjust the number of worker instances in response to the actual processing load. This prevents over-provisioning during low-throughput periods, directly reducing compute cost while maintaining performance during spikes.

Exam trap

Google Cloud often tests the misconception that preemptible VMs are always cost-effective for streaming workloads, but the trap here is that preemptible VMs are unsuitable for stateful streaming pipelines due to frequent preemption causing data reprocessing and instability.

377
MCQeasy

An application needs to store user profile data in a document database with flexible schema. The data is accessed frequently from a mobile app. Which Google Cloud database is BEST suited?

A.Cloud Bigtable
B.BigQuery
C.Cloud Spanner
D.Cloud Firestore
AnswerD

Cloud Firestore is a serverless document database with flexible schema and native mobile SDKs offering real-time sync and offline support. It suits frequently accessed user profile data from mobile apps, meeting the flexible-schema and mobile-access constraints.

Why this answer

Cloud Firestore is a NoSQL document database designed for mobile and web app development, offering flexible schema, real-time data synchronization, and automatic scaling. It directly supports frequent reads from mobile apps through its client SDKs and offline persistence, making it the best fit for storing user profile data with varying attributes.

Exam trap

The trap here is that candidates often confuse Cloud Firestore with Cloud Bigtable because both are NoSQL, but Bigtable lacks document flexibility, real-time sync, and mobile SDK support, which are essential for the described use case.

How to eliminate wrong answers

Option A is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput analytical and operational workloads (e.g., time-series, IoT), not for flexible document storage or mobile app real-time access. Option B is wrong because BigQuery is a serverless data warehouse for running SQL-based analytics on large datasets, not a transactional database for user profile reads/writes. Option C is wrong because Cloud Spanner is a globally distributed relational database with strong consistency and SQL support, but its rigid schema and higher latency for simple document operations make it overkill and less suitable for flexible schema mobile app data.

378
MCQmedium

A data team uses Cloud Dataproc to run nightly Spark jobs. The job volume has increased, and the cluster is often underutilized during the day. They want to reduce costs while ensuring jobs can scale when needed. Which strategy should they adopt?

A.Use preemptible workers for both primary and secondary nodes to minimize cost.
B.Manually scale the cluster up before nightly jobs and down after.
C.Use a cluster with a small number of primary workers and a large pool of preemptible workers, and enable autoscaling.
D.Use custom machine types with local SSDs for primary workers to improve I/O.
AnswerC

Autoscaling adds and removes worker nodes based on YARN pending-resource demand, so the cluster shrinks during idle daytime hours and expands for nightly Spark jobs. Preemptible workers cut compute cost for fault-tolerant batch work, while a small primary-worker core preserves HDFS durability. This directly satisfies the cost-reduction and elastic-scaling constraints.

Why this answer

It combines a small number of primary (non-preemptible) workers for reliability with a large pool of preemptible workers for cost-effective scaling, and enables autoscaling to dynamically adjust the cluster size based on workload. This minimizes cost during idle periods (preemptible instances are ~80% cheaper) while ensuring jobs can scale up quickly when needed, as autoscaling adds preemptible workers automatically. Preemptible workers are ideal for fault-tolerant Spark jobs that can handle node preemptions.

Exam trap

Google Cloud often tests the misconception that preemptible instances can be used for all nodes, but the trap here is that primary nodes require non-preemptible instances for cluster stability, while preemptible workers are only suitable for secondary (task) nodes in a fault-tolerant framework.

How to eliminate wrong answers

Option A is wrong because using preemptible workers for primary nodes is not allowed in Cloud Dataproc—primary nodes must be non-preemptible to ensure cluster stability and avoid data loss from coordinator failures. Option B is wrong because manual scaling is inefficient and error-prone for a nightly job pattern; it requires human intervention and cannot react to sudden workload spikes, leading to either underutilization or job delays. Option D is wrong because custom machine types with local SSDs improve I/O performance but do not address cost reduction or scaling needs; they increase cost without solving underutilization during the day.

379
MCQeasy

Which Google Cloud service is designed to replicate data from MySQL, PostgreSQL, and Oracle databases to BigQuery or Cloud Storage in near real-time?

A.Cloud Data Fusion
B.Datastream
C.Dataflow
D.Pub/Sub
AnswerB

Datastream is a serverless change data capture service that streams MySQL, PostgreSQL and Oracle changes into BigQuery or Cloud Storage with low latency. Its native source connectors and near real-time replication match the requirement, unlike batch-oriented transfer or replication tools.

Why this answer

Datastream is a serverless CDC service that ingests change data from relational databases into GCS or BigQuery.

380
MCQmedium

A mobile app uses Firestore to store user profiles. The app allows offline data creation and syncing when connectivity resumes. Which Firestore feature should the developer enable?

A.Set up a Firestore trigger to cache data in Cloud Memorystore
B.Enable offline persistence in the Firestore client SDK
C.Use Cloud Storage signed URLs for offline access
D.Enable Firestore multi-region replication
AnswerB

Enabling offline persistence in the Firestore client SDK caches data locally, so the app can create profiles without connectivity and automatically synchronise queued writes once the connection resumes, matching the offline creation and syncing scenario.

Why this answer

Firestore's offline persistence feature allows the client SDK to automatically cache data locally on the device. When the app creates or modifies data while offline, the SDK stores the changes in a local queue and syncs them with the Firestore backend once connectivity is restored. This is the correct and built-in mechanism for offline data creation and syncing.

Exam trap

Candidates often confuse Firestore's client-side offline persistence with server-side replication or caching features like multi-region replication or Cloud Memorystore, which are not designed for client-side offline data creation and syncing.

How to eliminate wrong answers

Option A is wrong because Cloud Memorystore is a managed Redis or Memcached service for caching in server-side applications, not a client-side offline cache; Firestore triggers are server-side functions that cannot cache data in Memorystore for offline client access. Option C is wrong because Cloud Storage signed URLs provide temporary, authenticated access to objects in Cloud Storage, not to Firestore documents, and they are used for online access, not offline data creation and syncing. Option D is wrong because multi-region replication improves availability and durability for Firestore databases but does not enable client-side offline caching or queuing of writes.

381
Multi-Selectmedium

A company wants to use Dataproc Metastore to manage metadata for their Spark jobs. Which TWO benefits does Dataproc Metastore provide?

Select 2 answers
A.Automatic scaling of compute resources
B.High availability with automatic failover
C.Fully managed Hive metastore service
D.Integration with BigQuery
E.Built-in data lineage tracking
AnswersB, C

Dataproc Metastore is a fully managed, highly available service that replicates metadata across zones and performs automatic failover without administrative intervention. This satisfies the stem's requirement for a managed metadata service for Spark jobs, eliminating the single point of failure inherent in a self-managed Hive Metastore deployment.

Why this answer

Option B is correct because Dataproc Metastore is a highly available, fully managed service that provides automatic failover across zones, ensuring metadata remains accessible even if a zone fails. Option C is correct because Dataproc Metastore is essentially a fully managed, serverless implementation of the Hive Metastore (HMS), which Spark and other engines use to store and retrieve metadata such as table schemas and partitions. Option A is incorrect because automatic scaling of compute resources is a feature of Dataproc clusters or autoscaling policies, not of the metadata service itself.

Option D is incorrect because Dataproc Metastore does not provide native BigQuery integration; it serves Hive/Spark-style metadata via the Hive Metastore Thrift protocol. Option E is incorrect because data lineage tracking is provided by tools like Data Catalog or Dataplex, not by Dataproc Metastore.

Exam trap

PDE often tests the distinction between Dataproc (compute) features like autoscaling and Dataproc Metastore (metadata) features like HA and managed Hive metastore, causing candidates to pick compute-related options.

382
MCQhard

You are designing a Dataflow pipeline that reads from a Cloud Storage bucket containing thousands of small JSON files, transforms the data, and writes to BigQuery. The pipeline is slow and expensive because of the large number of small files. You want to improve throughput and reduce cost. What should you do?

A.Use the FileIO transform with a match pattern and enable the `withHintMatchesManyFiles` option to optimize reading many small files.
B.Increase the number of Dataflow workers and set the autoscaling algorithm to THROUGHPUT_BASED to handle the small files in parallel.
C.Preprocess the data outside Dataflow to combine the small files into larger files (e.g., 100-500 MB each) in Cloud Storage, then run the Dataflow pipeline on the consolidated files.
D.Switch the pipeline to use the BigQuery Storage Read API to read the JSON files directly, bypassing Cloud Storage.
AnswerC

Consolidating small files into larger ones reduces the number of read operations and metadata overhead, which is the primary cause of slowness. Dataflow can then read each large file efficiently, improving throughput and lowering cost. This is the recommended pattern for many small files.

Why this answer

The bottleneck with thousands of small files is the per-file read and metadata overhead. Combining them into larger files before the Dataflow run reduces the number of reads and improves throughput. Increasing workers or using FileIO hints does not address the root cause, and the Storage Read API is for BigQuery tables, not Cloud Storage files.

Exam trap

The trap here is assuming that adding workers or using FileIO hints will fix small-file inefficiency, when the real fix is reducing the number of files by compaction.

383
MCQeasy

You need to track the lineage of data in BigQuery, showing how tables are derived from other tables via queries. Which service provides this capability?

A.BigQuery Lineage API
B.Cloud Composer
C.Cloud Data Catalog
D.Dataflow
AnswerA

The BigQuery Lineage API exposes column- and table-level lineage captured from query jobs, tracing how each table is derived from upstream tables. This directly satisfies the requirement to track derivation lineage, unlike Dataplex or Data Catalog, which handle discovery and governance metadata.

Why this answer

The BigQuery Lineage API provides programmatic access to lineage information, showing how tables are derived from other tables through queries, including column-level lineage. It captures dependencies from SQL queries, scheduled queries, and other BigQuery operations.

Exam trap

PDE often tests the confusion between Data Catalog (metadata discovery) and the Lineage API (derivation tracking) — candidates who equate 'catalog' with 'lineage' pick C.

How to eliminate wrong answers

Option B is wrong because Cloud Composer is a managed Apache Airflow service for workflow orchestration — it schedules jobs but does not track data lineage between BigQuery tables. Option C is wrong because Cloud Data Catalog (now Dataplex Universal Catalog) is a metadata management service for discovery and tagging, not a lineage tracking API for table derivation. Option D is wrong because Dataflow is a stream and batch processing service for building pipelines, not a lineage tracking system.

384
MCQhard

You have a BigQuery table `sales` with columns `order_id`, `customer_id`, `order_date` (DATE), and `amount` (NUMERIC). You need to create a new table that contains, for each customer, the total sales amount and the date of their most recent order. The result should include `customer_id`, `total_amount`, and `last_order_date`. Which SQL query achieves this correctly?

A.SELECT customer_id, SUM(amount) AS total_amount, order_date AS last_order_date FROM `project.dataset.sales` GROUP BY customer_id, order_date
B.SELECT customer_id, SUM(amount) AS total_amount, MAX(order_date) AS last_order_date FROM `project.dataset.sales` GROUP BY customer_id
C.SELECT customer_id, SUM(amount) AS total_amount, LAST_VALUE(order_date) AS last_order_date FROM `project.dataset.sales` GROUP BY customer_id
D.SELECT customer_id, SUM(amount) AS total_amount, MAX(order_date) OVER (PARTITION BY customer_id) AS last_order_date FROM `project.dataset.sales`
AnswerB

This query groups by `customer_id` and uses SUM to aggregate total sales and MAX to find the most recent order date. It correctly produces one row per customer with the desired columns. The GROUP BY clause ensures aggregation per customer, and the aggregate functions operate on the appropriate columns. This is the standard and efficient way to compute such per-customer metrics in BigQuery.

Why this answer

To compute per-customer aggregates, you use GROUP BY on `customer_id` and aggregate functions SUM and MAX. The SUM function totals the amounts, and MAX retrieves the latest order date. Other approaches using window functions or incorrect grouping do not produce the required single-row-per-customer result.

The correct query is straightforward and uses standard SQL aggregation.

Exam trap

The trap here is confusing window functions like LAST_VALUE with aggregate functions, or incorrectly grouping by order_date and producing multiple rows per customer.

385
MCQeasy

Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?

A.Dataflow
B.Dataproc
C.BigQuery
D.Cloud Dataprep
AnswerB

Dataproc is a managed Spark and Hadoop service, so existing Spark SQL jobs run with minimal refactoring. It satisfies the migration constraint by providing managed cluster provisioning and autoscaling while preserving the Spark SQL transformation logic used on-premises.

Why this answer

Dataproc is the managed Spark and Hadoop service on Google Cloud, purpose-built for running existing Spark SQL workloads with minimal changes. It allows you to spin up a cluster, run your Spark SQL transformations on large datasets stored in Cloud Storage or BigQuery, and then tear it down, making it the direct equivalent of an on-premises Hadoop cluster in the cloud.

Exam trap

Google Cloud certification exams often test the distinction between managed Spark (Dataproc) and serverless SQL (BigQuery) or Beam-based processing (Dataflow), trapping candidates who see 'SQL' and immediately think of BigQuery without recognizing the Spark SQL execution context.

How to eliminate wrong answers

Option A is wrong because Dataflow is a unified stream and batch processing service based on Apache Beam, not Spark SQL; migrating Spark SQL code to Dataflow would require rewriting the entire pipeline in Beam. Option C is wrong because BigQuery is a serverless data warehouse for SQL analytics, not a platform for running Spark SQL transformations; it does not support Spark execution engines. Option D is wrong because Cloud Dataprep is a visual data preparation tool for cleaning and structuring data, not a service for running Spark SQL jobs; it cannot execute custom Spark code.

386
MCQeasy

Which Dataflow feature allows you to package a pipeline into a reusable template that can be deployed with different parameters at runtime?

A.Cloud Dataproc
B.Classic Templates
C.Dataflow SQL
D.Flex Templates
AnswerD

Flex Templates package the pipeline as a Docker image plus a metadata file, allowing runtime parameterisation when launched via the gcloud CLI, REST API or Cloud Scheduler. This satisfies the requirement to reuse one pipeline definition across differing runtime parameters.

Why this answer

Dataflow Flex Templates allow you to containerize a pipeline and provide runtime parameters, enabling reusability across different environments or jobs.

387
MCQmedium

A data engineer needs to design a streaming pipeline that ingests events from multiple sources, enriches them with a lookup table stored in BigQuery (updated every hour), and writes the results to a BigQuery table for real-time dashboards. The pipeline must handle late-arriving data up to 1 hour. Which Dataflow feature should be configured to manage late data?

A.A custom watermark estimation function
B.Using side inputs with a periodic refresh
C.A trigger that fires on every late element
D.Allowed lateness on the window
AnswerD

Setting allowed lateness on the window lets Dataflow retain window state and emit updated results for up to one hour after the watermark passes, directly satisfying the stem's late-arriving data constraint. Unlike discarding mode, it triggers late firings so enriched events still reach the BigQuery dashboard table.

Why this answer

Allowed lateness on the window is the Dataflow feature that explicitly defines how long the pipeline will retain window state and accept late-arriving elements after the watermark has passed the end of the window. Setting allowed lateness to 1 hour lets Dataflow keep the window's state for that duration, so events arriving up to 60 minutes late are still processed and emitted with correct results. This directly satisfies the requirement to handle late data up to 1 hour without dropping or misassigning events.

Exam trap

PDE often tests the confusion between watermark estimation (which controls when the pipeline thinks data is complete) and allowed lateness (which controls how long late data is actually accepted), causing candidates to pick the watermark option.

How to eliminate wrong answers

Option A is wrong because a custom watermark estimation function only affects how the watermark advances (i.e., when the pipeline believes all data for a window has arrived); it does not retain window state or accept late elements after the watermark passes. Option B is wrong because side inputs with periodic refresh are used to enrich streaming elements with slowly changing reference data (like the BigQuery lookup table), not to handle late-arriving events. Option C is wrong because a trigger that fires on every late element only controls when results are emitted; it does not extend the window's lifetime, so late elements beyond the default allowed lateness would still be dropped.

388
MCQmedium

A company wants to orchestrate a multi-step data processing workflow that includes calling a Cloud Run service, waiting for its completion, and then running a BigQuery query. The workflow should be serverless and integrate with Cloud Events. Which Google Cloud service should they use?

A.Eventarc
B.Cloud Workflows
C.Cloud Composer
D.Cloud Dataflow
AnswerB

Cloud Workflows orchestrates the multi-step sequence, invoking the Cloud Run service, awaiting its completion, then running the BigQuery query, all serverless. It natively consumes Cloud Events via Eventarc triggers, satisfying the stem's event-integration constraint without managing infrastructure.

Why this answer

Cloud Workflows is the correct choice because it is a serverless workflow orchestrator that can coordinate multi-step processes involving Cloud Run and BigQuery. It natively supports waiting for asynchronous operations (like Cloud Run job completion) via its 'call' and 'wait' steps, and it can trigger subsequent steps such as BigQuery queries. Additionally, Cloud Workflows can be triggered by Cloud Events, making it fully integrated with the event-driven architecture described.

Exam trap

The trap here is that candidates confuse Google Cloud's Eventarc (event routing) with workflow orchestration, assuming that routing events alone can handle sequencing and waiting, when in fact Eventarc lacks the state management and step coordination required for multi-step workflows.

How to eliminate wrong answers

Option A is wrong because Eventarc is a service for routing events from various sources to targets (like Cloud Run, Cloud Functions), but it does not provide workflow orchestration capabilities such as waiting for completion or sequencing steps. Option C is wrong because Cloud Composer is a managed Apache Airflow service that is not serverless (it requires provisioning and managing a cluster of workers) and is overkill for a simple multi-step workflow; it is designed for complex, scheduled pipelines, not lightweight event-driven orchestration. Option D is wrong because Cloud Dataflow is a stream and batch data processing service (based on Apache Beam) that focuses on transforming data pipelines, not on orchestrating heterogeneous services like Cloud Run and BigQuery queries; it lacks native workflow sequencing and event-driven triggers.

389
MCQeasy

An engineer needs to create a Pub/Sub subscription that sends messages to an HTTPS endpoint. The endpoint must be able to acknowledge messages individually. Which type of subscription should they use?

A.Pull subscription
B.Push subscription
C.BigQuery subscription
D.Cloud Storage subscription
AnswerB

Push subscriptions deliver messages to an HTTPS endpoint and let the endpoint return an HTTP success code per message, which acknowledges that individual message. This satisfies the requirement for individual acknowledgement without the subscriber pulling.

Why this answer

Push subscriptions deliver messages to a configured HTTPS endpoint. The endpoint can acknowledge by returning a 200 status.

390
MCQhard

A data engineer is designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery using the Storage Write API in exactly-once mode. The pipeline must handle late-arriving data (up to 1 hour) and maintain correct aggregation results. Which trigger configuration should they use?

A.Default trigger without allowed lateness
B.After count trigger with count of 1000
C.Fixed time trigger every 5 minutes
D.After watermark trigger with allowed lateness of 1 hour
AnswerD

This configuration ensures late data within 1 hour is included in aggregations.

Why this answer

The After watermark trigger with allowed lateness of 1 hour is the correct choice because it emits results once the watermark passes the end of the window, while still accepting late-arriving events up to 1 hour past the watermark and including them in the correct event-time window. This aligns with the requirement to handle up to 1 hour of late data and maintain correct aggregation results. Note that exactly-once delivery to BigQuery is provided by the Storage Write API sink itself, not by the trigger; the trigger choice affects completeness and correctness of the aggregation, not the exactly-once guarantee.

Exam trap

Google often tests the misconception that any trigger with a time-based interval (like fixed 5-minute triggers) can handle late data, but only watermark-based triggers with allowed lateness correctly align with event-time windows and late-arriving records.

How to eliminate wrong answers

Option A is wrong because the default trigger fires only when the watermark passes, with no allowed lateness, so any data arriving more than a few seconds late would be dropped, violating the requirement to handle up to 1 hour of late data. Option B is wrong because an After count trigger of 1000 fires after every 1000 elements regardless of window boundaries or watermark, which does not respect the 1-hour lateness requirement and can cause incomplete or overlapping aggregations. Option C is wrong because a Fixed time trigger every 5 minutes fires periodically without regard to watermark or lateness, leading to premature emissions and incorrect aggregation results when late data arrives after the trigger fires.

391
MCQmedium

A retail company uses a machine learning model to predict inventory demand. The model is retrained weekly using Vertex AI Pipelines. Recently, the model's accuracy has degraded because the data distribution has shifted. Which action should you take to monitor and detect this drift automatically?

A.Enable Vertex AI Model Monitoring for the endpoint and configure alerting on feature drift
B.Set up alerts for when the model's mean absolute error exceeds a threshold on the evaluation dataset
C.Enable Cloud Logging for the prediction endpoint and search for error logs
D.Schedule a job to compare the distribution of incoming features with the training data using Cloud Dataflow
AnswerA

Vertex AI Model Monitoring on the endpoint computes feature distribution statistics against a training baseline and raises alerts when skew or drift exceeds configured thresholds. This detects the shifted data distribution automatically, satisfying the requirement to monitor and detect drift without manual inspection.

Why this answer

Vertex AI Model Monitoring is purpose-built to automatically detect feature drift and prediction drift on deployed endpoints. By enabling it and configuring alerting on feature drift, you can proactively identify when the distribution of incoming features deviates from the training data, which directly addresses the root cause of accuracy degradation without manual intervention.

Exam trap

Google Cloud often tests the distinction between monitoring model performance metrics (like MAE) versus monitoring input data distributions (feature drift), and candidates mistakenly choose a performance-based alerting option because they think accuracy degradation is the only signal, ignoring that drift detection is the proactive mechanism to catch the root cause before accuracy drops.

How to eliminate wrong answers

Option B is wrong because setting alerts on mean absolute error (MAE) on the evaluation dataset only detects performance degradation after the fact, not the underlying data distribution shift; it also requires ground truth labels, which may not be available in real time. Option C is wrong because Cloud Logging for the prediction endpoint captures request/response logs and error messages, but it does not perform statistical drift analysis or compare feature distributions. Option D is wrong because scheduling a job with Cloud Dataflow to compare distributions is a custom, manual approach that lacks the automated, integrated monitoring and alerting capabilities of Vertex AI Model Monitoring, and it introduces unnecessary operational overhead.

392
Multi-Selectmedium

Which THREE metrics should be monitored for a deployed machine learning model in production?

Select 3 answers
A.Number of replicas
B.Prediction error rate
C.Data drift detection
D.Training time
E.Prediction latency
AnswersB, C, E

Accuracy metric.

Why this answer

Prediction error rate (Option B) is a direct measure of model accuracy in production, reflecting how often the model's predictions deviate from actual outcomes. Monitoring this metric is essential for detecting model degradation, data quality issues, or concept drift that can silently reduce model performance over time.

Exam trap

Google Cloud often tests the distinction between operational metrics (like latency, error rate, drift) and development/infrastructure metrics (like training time, replica count) to see if candidates understand what is relevant for ongoing model monitoring versus model building or deployment scaling.

393
Multi-Selectmedium

A company is building a data pipeline that ingests streaming data from Pub/Sub, transforms it with Dataflow, and loads it into BigQuery. They want to handle malformed messages that cannot be parsed. Which TWO actions should they implement for error handling? (Choose 2)

Select 2 answers
A.Configure the pipeline to drop malformed messages silently
B.Use a side input to filter out malformed messages
C.Use a dead letter sink to write malformed messages to Cloud Storage or Pub/Sub for later analysis
D.Raise an exception in the DoFn to fail the pipeline immediately
E.Log the error and continue processing the next message
AnswersC, E

A dead letter sink captures unparseable messages instead of failing the pipeline, satisfying the requirement to handle malformed records. Dataflow's dead letter pattern writes them to Cloud Storage or Pub/Sub, preserving them for later analysis while the main pipeline continues.

Why this answer

Option C is correct because a dead letter sink (often implemented in Dataflow via a tagged output or a separate Pub/Sub topic/Cloud Storage path) captures unparseable records so they can be inspected and reprocessed without losing data or halting the pipeline. Option E is correct because logging the parsing error and continuing lets the pipeline keep processing valid messages, which is essential for a resilient streaming pipeline where one bad record should not stop the flow. Option A is not appropriate because silently dropping malformed messages loses data and provides no visibility for debugging or remediation.

Option B is not the right pattern because a side input is used to supply supplementary data to a DoFn, not to filter or quarantine malformed records. Option D is wrong because raising an exception in the DoFn would fail the pipeline immediately, causing unnecessary downtime and blocking valid messages from being processed.

Exam trap

Google Cloud often tests the misconception that raising an exception (Option D) is acceptable for error handling in streaming pipelines, but the correct approach is to isolate failures using a dead letter sink (Option C) while logging errors (Option E) to maintain pipeline continuity.

394
Multi-Selecthard

You are designing a data processing architecture on Google Cloud. You need to ingest data from multiple sources, including streaming events and batch files, and process them to produce a unified dataset for analytics. The solution must support both real-time and historical processing with the same codebase, and be able to handle late-arriving data. Which two Google Cloud services should you use together to achieve this? (Choose two.)

Select 2 answers
A.Cloud Composer
B.Cloud Pub/Sub
C.Cloud Dataflow
D.Cloud Dataproc
E.BigQuery
AnswersB, C

Cloud Pub/Sub is a scalable, serverless messaging service for ingesting streaming data. It can handle high-throughput event streams and integrate seamlessly with Dataflow. Pub/Sub ensures reliable delivery and can buffer messages, which is essential for ingesting streaming events into a unified pipeline that also processes batch data.

Why this answer

To achieve unified batch and stream processing with the same codebase and handle late data, you need a processing engine that supports Apache Beam, such as Cloud Dataflow, and a scalable ingestion service for streaming data, such as Cloud Pub/Sub. Dataflow's Beam model allows you to write one pipeline that works for both batch and streaming, with built-in support for event-time windowing and late data. Pub/Sub provides reliable, scalable ingestion for streaming events.

Exam trap

The trap here is assuming that any data processing service can handle both batch and streaming with the same code, or that orchestration services like Cloud Composer can replace a stream processing engine.

395
MCQeasy

You need to copy a 3 TB dataset from an on-premises Hadoop Distributed File System (HDFS) cluster to Cloud Storage over a 1 Gbps dedicated interconnect. The data is a one-time historical backfill and must be transferred with integrity verification. You want to minimize operational overhead and avoid writing custom transfer code. Which service should you use?

A.Cloud Data Fusion with an HDFS source plugin
B.gcloud storage cp with parallel composite uploads
C.DistCp job writing directly to a Cloud Storage bucket
D.Storage Transfer Service with an HDFS source
AnswerD

Storage Transfer Service natively supports on-premises HDFS as a source and Cloud Storage as a destination. It performs parallel transfers, integrity checks, and retries without custom code, which fits the one-time historical backfill over a dedicated interconnect. It also handles the large file set efficiently and provides transfer logs and monitoring out of the box.

Why this answer

Storage Transfer Service is purpose-built for moving large datasets from on-premises sources including HDFS into Cloud Storage. It handles parallelism, retries, integrity verification, and monitoring without custom code, which matches the requirement to minimize operational overhead for a one-time backfill over a dedicated interconnect.

Exam trap

The trap here is assuming that any tool that can read HDFS is equally suited for bulk transfer, when transfer-specific managed services provide integrity and retry behavior that pipeline tools do not.

396
MCQmedium

You need to inspect a BigQuery table for sensitive data such as credit card numbers and apply masking. Which GCP service should you use to identify and de-identify the data?

A.IAM Recommender
B.Cloud KMS
C.Dataplex
D.Cloud Data Loss Prevention (DLP)
AnswerD

Cloud Data Loss Prevention scans BigQuery tables using built-in infoTypes to detect credit card numbers, then applies de-identification transforms such as masking, tokenisation or redaction directly on the findings. This satisfies the stem's requirement to both identify sensitive data and mask it within BigQuery.

Why this answer

Cloud Data Loss Prevention (DLP) is Google Cloud's service for discovering, classifying, and de-identifying sensitive data such as credit card numbers, SSNs, and emails. It provides infoType detectors (including CREDIT_CARD_NUMBER) and de-identification transforms like masking, tokenization, and format-preserving encryption. It integrates directly with BigQuery for scanning tables and applying masking at query or storage time.

Exam trap

PDE often tests the confusion between governance/orchestration services (Dataplex) and the actual detection-and-de-identification engine (Cloud DLP), so candidates pick Dataplex when the task is specifically identifying and masking sensitive data.

How to eliminate wrong answers

Option A is wrong because IAM Recommender suggests least-privilege role changes based on usage — it has no data inspection or de-identification capability. Option B is wrong because Cloud KMS manages encryption keys for data at rest; it encrypts but does not identify sensitive content or apply field-level masking. Option C is wrong because Dataplex is a data governance and lakehouse management service that can orchestrate discovery and lineage, but the actual sensitive-data detection and de-identification engine it invokes is Cloud DLP, not Dataplex itself.

397
MCQmedium

A media company uses Cloud Data Loss Prevention (DLP) API to inspect and de-identify sensitive data before loading into BigQuery. They want to reduce costs by sampling the data during inspection. Which configuration should they use?

A.Use the 'ROWS' limit in the inspection job.
B.Set the sample method to 'RANDOM' with a percentage.
C.Use a hybrid inspection with a BigQuery sample table.
D.Use the 'BYTES_LIMIT' parameter.
AnswerB

The RANDOM sample method with a percentage inspects only a statistical subset of records, reducing both the volume of data processed and the associated DLP API charges. This satisfies the stated cost-reduction goal while still providing representative inspection coverage before loading into BigQuery.

Why this answer

The Cloud DLP API supports a 'sample_method' of 'RANDOM' with a 'sampling_percentage' to inspect only a random subset of rows. This directly reduces the volume of data scanned, lowering costs while still providing statistically representative coverage for sensitive data discovery.

Exam trap

The trap here is that candidates confuse 'limiting rows/bytes' (which scans sequentially from the start) with 'random sampling' (which distributes inspection across the entire dataset), leading them to pick options A or D, which do not achieve representative cost reduction.

How to eliminate wrong answers

Option A is wrong because the 'ROWS' limit in an inspection job caps the total number of rows scanned but does not sample randomly; it stops after scanning that many rows from the start, which can miss sensitive data in later rows and does not provide representative sampling. Option C is wrong because hybrid inspection with a BigQuery sample table requires manually creating and maintaining a separate table, adding complexity and storage costs, whereas the DLP API's built-in sampling is simpler and directly integrated. Option D is wrong because 'BYTES_LIMIT' limits the total bytes scanned but, like 'ROWS', scans sequentially from the beginning and does not perform random sampling, leading to biased results and potential cost inefficiency.

398
MCQmedium

Refer to the exhibit. A BigQuery dataset has the IAM policy shown above. An analyst is trying to run a SELECT query on a table in this dataset but receives an 'Access Denied' error. What is the most likely reason?

A.The analyst does not have permission to list datasets in the project.
B.The analyst only has the roles/bigquery.metadataviewer role, which does not allow reading table data.
C.The table is in a different region than the dataset, and the analyst's query is not cross-region compatible.
D.The analyst has not been granted the 'bigquery.jobs.create' permission to run queries.
AnswerB

The roles/bigquery.metadataviewer role grants read access to dataset and table metadata only, not to the underlying rows. Running SELECT requires roles/bigquery.dataViewer, so the analyst can list the table but cannot read its data, producing the Access Denied error.

Why this answer

The roles/bigquery.metadataviewer role grants permissions to view table and dataset metadata (e.g., table names, schemas) but does not include the bigquery.tables.getData permission required to read table rows. Therefore, when the analyst runs a SELECT query, BigQuery denies access because the role lacks the data-reading privilege. This is the most likely reason for the 'Access Denied' error.

Exam trap

Google Cloud often tests the distinction between metadata-viewing roles and data-reading roles, trapping candidates who assume that being able to see table names and schemas implies permission to query the data.

How to eliminate wrong answers

Option A is wrong because listing datasets is not required to run a SELECT query; the error is about reading table data, not dataset enumeration. Option C is wrong because BigQuery does not enforce cross-region compatibility at the dataset-table level; tables reside within the same dataset and region, and cross-region queries are allowed with appropriate permissions. Option D is wrong because the 'bigquery.jobs.create' permission is needed to submit a query job, but the error specifically indicates a data access issue, not a job creation failure; the analyst likely has this permission if they can attempt a query.

399
Multi-Selecthard

A data pipeline reads thousands of JSON files from Cloud Storage, processes them with Cloud Dataflow, and writes to BigQuery. The pipeline sometimes fails because of malformed JSON records. Which three steps should the data engineering team take to improve pipeline reliability? (Choose THREE.)

Select 3 answers
A.Integrate Cloud Pub/Sub as an intermediary to buffer and allow message retry
B.Use a try-catch block in the pipeline to retry processing failed records
C.Create a Cloud Monitoring alert on pipeline failures
D.Add schema validation before processing to reject invalid JSON records
E.Implement a dead-letter queue in the Dataflow pipeline to store failed records for later analysis
AnswersA, D, E

Pub/Sub can retry delivery of messages, improving reliability.

Why this answer

Integrating Cloud Pub/Sub as an intermediary decouples the ingestion of JSON files from the Dataflow pipeline. Pub/Sub provides at-least-once delivery and automatic retries for messages that are not acknowledged, which buffers against transient failures and malformed records. This allows the pipeline to pull messages at its own pace and retry processing without losing data.

Exam trap

The trap here is that candidates often confuse reactive monitoring (Option C) with proactive reliability improvements, or they assume a simple try-catch block (Option B) is sufficient in a distributed processing framework like Dataflow, where fault tolerance requires persistent retry mechanisms and dead-letter queues.

400
MCQeasy

A team needs to migrate an existing on-premises Hadoop Hive workload to Google Cloud. They want to minimize code changes and use a managed service for transient clusters. Which service should they choose?

A.Cloud Dataflow
B.Cloud Dataprep
C.Cloud Dataproc
D.BigQuery
AnswerC

Cloud Dataproc runs Apache Hive natively on managed, ephemeral clusters, so existing HiveQL scripts migrate with minimal rewriting. It satisfies both constraints: the managed service removes cluster administration, and transient clusters spin up only for job duration, cutting cost versus persistent on-premises infrastructure.

Why this answer

Cloud Dataproc is the correct choice because it is a managed Spark and Hadoop service that supports Hive workloads natively, allowing you to run existing Hive scripts with minimal changes. It also supports transient clusters, which can be automatically scaled up and down, aligning with the requirement for transient clusters.

Exam trap

The trap here is that candidates often confuse Cloud Dataflow's ability to process batch data with Hadoop compatibility, but Dataflow does not support Hive or transient Hadoop clusters, making Dataproc the only correct option for minimizing code changes.

How to eliminate wrong answers

Option A is wrong because Cloud Dataflow is a unified stream and batch data processing service based on Apache Beam, not designed for Hive workloads or transient Hadoop clusters. Option B is wrong because Cloud Dataprep is a data preparation and cleaning service (based on Trifacta) that does not run Hive or provide transient clusters. Option D is wrong because BigQuery is a serverless data warehouse that does not support Hive execution engines or transient clusters; migrating Hive to BigQuery would require significant code changes.

401
Multi-Selectmedium

A data engineer is monitoring a Dataflow streaming pipeline and notices that the 'System Lag' metric is increasing. Which TWO actions should be taken to diagnose the issue?

Select 2 answers
A.Check the Dataflow monitoring UI for each stage's throughput and backlog.
B.Cancel the pipeline and restart with a larger initial worker count.
C.Increase the maximum number of workers to handle backlog.
D.Examine the worker logs for error messages or stack traces.
E.Increase the BigQuery quota for streaming inserts.
AnswersA, D

Identifies bottleneck stages.

Why this answer

The Dataflow monitoring UI provides per-stage metrics such as throughput and backlog, which directly indicate where data is accumulating. By examining these metrics, you can identify the specific stage causing the increasing system lag, enabling targeted troubleshooting without unnecessary pipeline changes.

Exam trap

Google Cloud often tests the distinction between diagnostic actions and remedial actions; the trap here is that candidates confuse scaling up workers (a fix) with diagnosing the root cause of the lag.

402
MCQmedium

You deployed a model on Vertex AI Endpoints using a custom container. The model serves predictions but the latency is higher than expected. You suspect the container is not making full use of the CPU resources. What should you do to reduce latency?

A.Modify the container to use multi-threading or increase the number of workers in the prediction server (e.g., Gunicorn workers).
B.Enable response caching on the endpoint.
C.Change the machine type to a GPU-accelerated machine.
D.Increase the number of nodes by adjusting autoscaling limits.
AnswerA

A single-threaded prediction server serialises requests, leaving CPU cores idle and inflating latency. Enabling multi-threading or raising Gunicorn worker counts lets the container process concurrent inferences in parallel, directly addressing the underused CPU identified in the stem.

Why this answer

High latency in a CPU-based custom container often stems from underutilizing available CPU cores. By increasing the number of workers (e.g., Gunicorn workers) or enabling multi-threading, you allow the prediction server to handle multiple requests concurrently, reducing queue time and improving throughput. This directly addresses the symptom of the container not making full use of CPU resources.

Exam trap

Google Cloud often tests the misconception that scaling out (adding more nodes) or upgrading hardware (GPU) is the default fix for latency, when the real issue is often software-level concurrency configuration within the container.

How to eliminate wrong answers

Option B is wrong because response caching reduces latency only for repeated identical requests, not for the general case of underutilized CPU resources; it does not improve concurrent request handling. Option C is wrong because switching to a GPU-accelerated machine would only help if the model benefits from GPU parallelism (e.g., deep learning models), but the question states the container is not making full use of CPU resources, implying the bottleneck is software configuration, not hardware type. Option D is wrong because increasing the number of nodes via autoscaling adds more instances but does not fix the per-instance CPU underutilization; it may even increase cost without addressing the root cause of inefficient request handling within each container.

403
MCQmedium

A logistics company uses Cloud Pub/Sub to ingest shipment tracking events. They want to archive all events to Cloud Storage for long-term retention and also process them in real time with Dataflow. The events are published to a single topic. Which design should the data engineer use to ensure both archiving and real-time processing without data loss?

A.Create two subscriptions on the topic: one for the Dataflow pipeline and one for a Cloud Function that writes to Cloud Storage.
B.Configure the Pub/Sub topic to write to Cloud Storage directly using a Pub/Sub to Cloud Storage subscription.
C.Use a single subscription for both the Dataflow pipeline and a Cloud Function that writes to Cloud Storage.
D.Use a single subscription for the Dataflow pipeline, and have the pipeline write to both Cloud Storage and its real-time processing output.
AnswerA

Pub/Sub allows multiple subscriptions on a single topic, each receiving a copy of every message. By creating one subscription for Dataflow and another for a Cloud Function that archives to Cloud Storage, both consumers independently receive all events. This ensures no data loss and allows independent processing. This is the standard fan-out pattern.

Why this answer

The correct design is to create two separate subscriptions on the Pub/Sub topic: one for Dataflow and one for a Cloud Function that archives to Cloud Storage. This fan-out pattern ensures that each event is delivered to both consumers independently, preventing data loss and allowing each to process at its own pace. It decouples archiving from real-time processing.

Exam trap

The trap here is assuming that a single subscription can be shared by multiple consumers to receive all messages, but Pub/Sub delivers each message to only one consumer per subscription.

404
MCQmedium

A company uses Cloud Dataproc to run Spark ML training jobs. They want to persist the trained models and metadata in a Hive-compatible metastore. Which Dataproc feature should they use?

A.Cloud Hive Metastore (self-managed)
B.Cloud Bigtable
C.Dataproc Metastore
D.Cloud Data Catalog
AnswerC

Dataproc Metastore provides a fully managed, Hive-compatible metastore service that persists table metadata and model artefacts independently of cluster lifecycle. This satisfies the requirement to retain trained models and metadata in a Hive-compatible store across ephemeral Dataproc clusters.

Why this answer

Dataproc Metastore is a fully managed, Hive-compatible metastore service on Google Cloud that integrates natively with Dataproc clusters, allowing Spark and Hive jobs to persist and share table metadata and schemas. It provides a serverless, scalable alternative to running a self-managed Hive metastore on a cluster, and it supports the Hive Metastore Thrift API so existing Spark ML workflows can store model metadata without code changes. This directly meets the requirement for a Hive-compatible metastore for trained models and metadata.

Exam trap

PDE often tests the confusion between metadata storage services (Dataproc Metastore vs. Data Catalog) and database services (Bigtable), causing candidates to choose a non-Hive-compatible option for metastore requirements.

How to eliminate wrong answers

Option A is wrong because a self-managed Cloud Hive Metastore requires provisioning and operating a Hive metastore on a VM or cluster, adding operational overhead and not being a managed Dataproc feature. Option B is wrong because Cloud Bigtable is a NoSQL wide-column database for low-latency workloads, not a Hive-compatible metadata store. Option D is wrong because Cloud Data Catalog is a metadata management and discovery service, not a Hive Metastore replacement that Spark/Hive can read and write table metadata to.

405
MCQmedium

A company is using Cloud Storage to store raw logs. They want to use Cloud Data Fusion to transform and load the data into BigQuery on a daily schedule. The transformations are complex and involve joining multiple datasets. What is the most efficient way to run these pipelines?

A.Use Cloud Composer to orchestrate Dataproc jobs that run the transformations
B.Use Cloud Functions to trigger a Dataflow job that does the transformations
C.Use Cloud Data Fusion to design the pipeline and schedule it to run on a Dataproc cluster
D.Use Cloud Dataprep to design the transformation and export to BigQuery
AnswerC

Cloud Data Fusion pipelines execute on ephemeral Dataproc clusters, which provide the distributed compute needed for complex joins across multiple datasets. Scheduling daily runs on Dataproc satisfies the stem's transformation complexity while remaining cost-efficient between executions.

Why this answer

Cloud Data Fusion is a fully managed, code-free ETL/ELT service built on CDAP that provides a visual pipeline designer with a rich set of plugins (including BigQuery sinks) and can execute pipelines on ephemeral Dataproc clusters. For complex transformations involving joins across multiple datasets, Data Fusion's Wrangler and Joiner transforms plus its scheduling capability make it the most efficient native choice. Scheduling the pipeline to run on a Dataproc cluster gives the compute needed for joins while keeping orchestration managed.

Exam trap

PDE often tests whether candidates confuse Cloud Data Fusion (visual ETL with scheduling) with Cloud Dataprep (data preparation only) or Cloud Composer (general orchestration) — the key differentiator is that Data Fusion is the managed ETL tool that natively schedules and runs on Dataproc.

How to eliminate wrong answers

Option A is wrong because Cloud Composer orchestrating Dataproc jobs requires writing Spark code and DAGs manually — it is not the most efficient way when Data Fusion already provides the visual transform and scheduling layer. Option B is wrong because Cloud Functions is serverless and short-lived, unsuitable for orchestrating complex multi-dataset joins, and Dataflow would require custom Apache Beam code rather than the requested Data Fusion approach. Option D is wrong because Cloud Dataprep is a data-preparation tool for profiling and cleaning, not a full ETL orchestrator with scheduling and complex join semantics at scale.

406
Multi-Selecthard

A payment processing company needs to detect fraudulent transactions in real time. The system must have sub-second latency for high-value transactions and use a machine learning model. Which two components should be part of the architecture? (Choose TWO.)

Select 2 answers
A.Cloud Storage for transaction logs
B.Bigtable to store user profiles and transaction history for fast lookups
C.Dataflow for stream processing with sliding windows
D.Cloud SQL to store reference data
E.Cloud Functions for long-running batch model training
AnswersB, C

Bigtable offers sub-millisecond latency for point lookups, essential for real-time fraud scoring.

Why this answer

Bigtable is a fully managed, scalable NoSQL database that provides consistent sub-10ms latency for high-throughput read/write operations, making it ideal for real-time lookups of user profiles and transaction history in fraud detection. Its ability to handle large volumes of data with low latency supports the sub-second requirement for high-value transactions.

Exam trap

Google Cloud often tests the distinction between storage services optimized for real-time access (Bigtable) versus batch/archive (Cloud Storage) and between stream processing (Dataflow) versus batch processing or short-lived compute (Cloud Functions).

407
MCQeasy

You are ingesting streaming data into BigQuery using the Storage Write API. The data arrives with occasional duplicate records due to at-least-once delivery semantics from the source system. You need to ensure that the final BigQuery table contains no duplicates based on a unique event_id column. Which approach should you use?

A.Use the Storage Write API with application-created streams and specify a unique offset for each record to enable exactly-once semantics.
B.Enable the exactly-once delivery semantics in the Storage Write API by using the default stream and specifying a unique offset for each record.
C.Use BigQuery's MERGE statement to upsert records into the table based on event_id, running it periodically to deduplicate.
D.Create a unique constraint on the event_id column in the BigQuery table to automatically reject duplicates.
AnswerA

The Storage Write API supports exactly-once semantics when using application-created streams. By assigning a unique offset to each record within a stream, BigQuery can deduplicate records on write. This ensures that even if the client retries, duplicates are not inserted. This is the recommended approach for streaming ingestion with deduplication requirements. It provides low-latency, exactly-once delivery without additional batch processing.

Why this answer

To achieve exactly-once semantics with the Storage Write API, you must use application-created streams and assign a unique offset to each record. This allows BigQuery to deduplicate records on write, ensuring no duplicates even with retries. The default stream does not provide exactly-once semantics, and BigQuery lacks unique constraints.

MERGE is a batch alternative but not ideal for streaming.

Exam trap

The trap here is assuming that the Storage Write API's default stream or a unique constraint can provide deduplication, when in fact exactly-once requires application-created streams with offsets.

408
MCQeasy

A startup wants to build a data lake on Google Cloud using Cloud Storage. They need to store raw data in its original format for future analysis. Which storage class should they use to optimize for cost given that data will be accessed occasionally after the first month?

A.Nearline storage class
B.Coldline storage class
C.Standard storage class
D.Archive storage class
AnswerA

Nearline storage suits data accessed less than once per month, offering lower storage cost than Standard while charging retrieval fees. Raw data retained in original format for occasional future analysis after the first month matches Nearline's access pattern and cost profile.

Why this answer

Nearline storage class is the optimal choice because it offers low-cost storage for data accessed less than once a month, with a 30-day minimum storage duration. Since the data is accessed occasionally after the first month, Nearline provides significant cost savings over Standard while still offering low-latency access (milliseconds) suitable for analytics. Coldline and Archive have lower storage costs but impose higher retrieval fees and minimum storage durations (90 and 365 days respectively), making them more expensive for data that is accessed occasionally within the first year.

Exam trap

Google Cloud often tests the misconception that lower storage cost always means lower total cost, ignoring the impact of retrieval fees and minimum storage duration penalties, which can make Coldline or Archive more expensive for data accessed occasionally within the first year.

How to eliminate wrong answers

Option B (Coldline) is wrong because it is designed for data accessed less than once a quarter (90-day minimum storage duration) and has higher retrieval costs, making it more expensive than Nearline for data accessed occasionally after the first month. Option C (Standard) is wrong because it is optimized for frequently accessed data (no minimum storage duration) and has the highest storage cost, which is not cost-effective for data that is only accessed occasionally. Option D (Archive) is wrong because it is intended for long-term archival data accessed less than once a year (365-day minimum storage duration) and has very high retrieval costs and latency (hours), making it unsuitable for occasional access within a year.

409
MCQmedium

A data engineer is designing a batch data pipeline that reads Avro files from Cloud Storage, transforms data using Apache Beam, and writes to BigQuery. The pipeline must handle daily runs and backfills. Which runner should they use?

A.FlinkRunner
B.DataflowRunner
C.SparkRunner
D.DirectRunner
AnswerB

DataflowRunner executes Apache Beam pipelines on Google Cloud, reading from Cloud Storage and writing to BigQuery with autoscaling. It handles both scheduled daily runs and historical backfills by replaying the same pipeline over specified date ranges without infrastructure management.

Why this answer

DataflowRunner is the correct choice because it is the fully managed service runner for Apache Beam on Google Cloud, optimized for batch and streaming pipelines. It automatically handles scaling, resource management, and exactly-once processing semantics, which are essential for reliable daily runs and backfills with Avro files from Cloud Storage and BigQuery sinks.

Exam trap

The trap here is that candidates may confuse the runner with the execution engine, assuming that any distributed runner (Flink, Spark) is suitable for production, when the question specifically tests knowledge of Google Cloud-native services and the need for managed infrastructure for batch pipelines with backfills.

How to eliminate wrong answers

Option A is wrong because FlinkRunner is designed for running Beam pipelines on Apache Flink clusters, which require manual cluster management and are not natively integrated with Google Cloud services like Cloud Storage and BigQuery. Option C is wrong because SparkRunner runs Beam pipelines on Apache Spark, which is not a managed service on Google Cloud and lacks the seamless integration with Cloud Storage and BigQuery that DataflowRunner provides. Option D is wrong because DirectRunner is intended for local testing and development only, not for production workloads or handling large-scale daily runs and backfills.

410
Multi-Selectmedium

Which TWO are best practices for managing a Cloud Dataflow pipeline in production?

Select 2 answers
A.Always use batch mode for streaming data to reduce cost
B.Disable autoscaling to keep compute costs predictable
C.Set up Cloud Monitoring alerts based on Dataflow job metrics
D.Use pipeline updates (update) to modify running streaming pipelines
E.Restart the pipeline when code changes are needed
AnswersC, D

Cloud Monitoring alerts on Dataflow job metrics such as system lag, backlog and worker errors surface degradation before pipelines fail, enabling proactive intervention. This is a standard production operational practise for streaming and batch jobs.

Why this answer

Option C is correct because Cloud Monitoring integrates with Dataflow job metrics (such as system lag, data watermark, and element count) so you can create alerting policies that detect stuck pipelines, backlog growth, or failures and respond before SLAs are breached. Option D is correct because Dataflow supports updating a running streaming pipeline via the --update and --transformNameMapping flags, which lets you change code or adjust transforms while preserving the pipeline's state and avoiding downtime. Option A is wrong because batch mode cannot process unbounded streaming data; streaming pipelines require streaming mode, and switching to batch would break the use case rather than reduce cost.

Option B is wrong because disabling autoscaling removes Dataflow's ability to dynamically adjust worker count to workload, hurting performance and often increasing cost due to over- or under-provisioning. Option E is wrong because restarting a streaming pipeline loses in-flight state and causes downtime; the recommended approach is an update rather than a stop-and-restart.

Exam trap

Google Cloud often tests the misconception that disabling autoscaling or restarting pipelines is acceptable for cost control or simplicity, when in fact these actions violate production best practices for reliability and data integrity.

411
MCQhard

A financial services company uses Cloud Pub/Sub with ordering keys to process transactions in order. Some messages are failing processing and getting stuck. The team wants to ensure that if a message fails, it can be reprocessed later without blocking subsequent messages. What should they implement?

A.Create multiple subscriptions for the same topic
B.Use a pull subscription with flow control settings
C.Configure a dead letter topic and handle the failed message separately
D.Increase the acknowledgment deadline to 600 seconds
AnswerC

A dead letter topic captures messages that exceed the maximum delivery attempts, removing them from the subscription so they stop blocking the ordered backlog. The failed transaction is then reprocessed separately, preserving ordering for subsequent messages.

Why this answer

A dead letter topic (DLT) allows failed messages to be moved aside after exhausting retry attempts, so they do not block the processing of subsequent ordered messages. In Cloud Pub/Sub, ordering keys require messages with the same key to be delivered in order; if a message fails and is not acknowledged, it blocks all later messages with the same key. By configuring a dead letter topic, the failed message is automatically forwarded to the DLT after a maximum of 5 delivery attempts (default), and the original subscription can continue processing the next messages in order.

The team can then reprocess the failed message from the DLT separately, without affecting the order of other messages.

Exam trap

Google Cloud often tests the misconception that increasing the acknowledgment deadline or adding flow control can resolve stuck messages with ordering keys, but the real solution is to use a dead letter topic to offload the failing message and unblock the ordered stream.

How to eliminate wrong answers

Option A is wrong because creating multiple subscriptions for the same topic does not solve the blocking issue; each subscription independently receives all messages, but within a single subscription, ordering keys still cause a failed message to block subsequent messages with the same key. Option B is wrong because pull subscriptions with flow control settings only limit the rate of message delivery and do not handle failed messages that are stuck; they do not provide a mechanism to move failed messages out of the way to unblock ordering. Option D is wrong because increasing the acknowledgment deadline to 600 seconds only gives the subscriber more time to process a message before it is redelivered, but it does not prevent a persistently failing message from blocking subsequent ordered messages indefinitely.

412
MCQeasy

A data engineer needs to process data in a Dataflow pipeline that reads from a Pub/Sub topic. The pipeline must group events into 5-minute windows and compute the average value per key. Which Beam transform should they use after windowing?

A.Combine.perKey
B.ParDo
C.GroupByKey
D.CoGroupByKey
AnswerA

Combine.perKey performs per-key aggregation after windowing, computing the average value for each key within each 5-minute window. It satisfies the requirement to group events and average per key, unlike global combines that would merge across keys.

Why this answer

Combine.perKey is the correct transform because it performs a per-key aggregation (here, computing the average value) after windowing, combining elements within each key and window efficiently. It is a fused Combine operation that reduces data before shuffling, making it more efficient than GroupByKey followed by a separate aggregation. Since the requirement is to compute an average per key within 5-minute windows, Combine.perKey directly expresses that intent.

Exam trap

The trap is choosing GroupByKey because it 'groups by key,' but the question asks for an aggregation — Combine.perKey is the efficient, purpose-built transform for per-key aggregation and avoids the full shuffle penalty of GroupByKey.

How to eliminate wrong answers

Option B is wrong because ParDo is a general-purpose element-wise transform for mapping/filtering individual elements; it does not perform per-key aggregation across a window on its own. Option C is wrong because GroupByKey groups all values for a key into an iterable but does not compute an aggregate — you would still need a subsequent transform, and it shuffles all values without pre-aggregation, which is less efficient. Option D is wrong because CoGroupByKey joins multiple PCollections by key, which is unnecessary here since only one source (Pub/Sub) is being aggregated.

413
MCQeasy

A startup is building a data lake on Google Cloud. They need to store raw JSON logs in a cost-effective manner and later query them using SQL with minimal transformation. The logs are infrequently accessed but must be retained for 7 years for compliance. Which storage solution should they use?

A.Cloud Bigtable with a column family for JSON logs and a HBase client for querying.
B.Cloud SQL for PostgreSQL with JSONB columns and scheduled exports to Cloud Storage.
C.Cloud Storage with Archive storage class and BigQuery external tables.
D.Cloud Storage with Nearline storage class and BigQuery external tables.
AnswerC

Archive storage is the lowest-cost option for long-term retention and is ideal for data accessed less than once a year. BigQuery external tables allow querying JSON data directly from Cloud Storage without loading, satisfying the SQL query requirement with minimal transformation. This combination is both cost-effective and functional for infrequent access over 7 years.

Why this answer

Archive storage in Cloud Storage is the most cost-effective for long-term retention of infrequently accessed data. BigQuery external tables enable SQL querying of JSON logs directly from Cloud Storage without loading or transformation. This meets the cost, retention, and query requirements.

The other options use higher-cost storage or databases not optimized for this use case.

Exam trap

The trap here is choosing Nearline or Coldline for 7-year retention when Archive is specifically designed for the lowest-cost, long-term storage.

414
MCQmedium

You need to store and query a large dataset of customer profiles. The data is semi-structured and frequently updated. The application requires offline support for mobile users. Which database is MOST appropriate?

A.Firestore
B.BigQuery
C.Cloud Bigtable
D.Cloud SQL
AnswerA

Firestore stores semi-structured documents with flexible schemas and offers real-time synchronisation plus offline persistence on mobile clients, so users can query and update profiles without connectivity. This directly satisfies the stem's offline support requirement, unlike relational databases that demand fixed schemas and constant connectivity.

Why this answer

Firestore is the most appropriate choice because it is a NoSQL document database designed for semi-structured data, real-time synchronization, and offline support. It provides built-in offline persistence for mobile clients, allowing users to read and write data even without network connectivity, and automatically syncs changes when the connection is restored. This directly meets the requirements of semi-structured data, frequent updates, and offline mobile support.

Exam trap

Google Cloud often tests the distinction between databases designed for transactional/operational workloads (like Firestore) versus analytical/warehouse databases (like BigQuery), and the trap here is assuming that any NoSQL database (like Bigtable) supports offline mobile sync, when in fact only Firestore provides native offline persistence and real-time synchronization for mobile clients.

How to eliminate wrong answers

Option B (BigQuery) is wrong because it is a serverless data warehouse optimized for analytical queries on large datasets, not for transactional or real-time updates, and it lacks native offline support for mobile applications. Option C (Cloud Bigtable) is wrong because it is a wide-column NoSQL database designed for high-throughput, low-latency workloads like time-series or IoT data, but it does not support offline mobile synchronization or semi-structured document models. Option D (Cloud SQL) is wrong because it is a relational database (MySQL, PostgreSQL, SQL Server) requiring a fixed schema, which is unsuitable for semi-structured data, and it does not provide built-in offline support for mobile clients.

415
MCQeasy

Which Google Cloud service provides a visual interface for building ETL pipelines using a drag-and-drop design and includes pre-built transforms from a marketplace?

A.Dataproc
B.Cloud Data Fusion
C.Dataprep
D.BigQuery
AnswerB

Cloud Data Fusion provides a graphical drag-and-drop interface for building ETL pipelines, with a reusable plugin marketplace of pre-built transforms and connectors. This directly satisfies the stem's requirement for visual pipeline design plus marketplace transforms, unlike code-first services such as Dataflow or Dataproc.

Why this answer

Cloud Data Fusion is Google Cloud's fully managed, code-free ETL/ELT service built on the open-source CDAP project, offering a graphical drag-and-drop pipeline designer and a marketplace of pre-built plugins and transformations. It lets data engineers build batch and streaming pipelines visually and deploy them to ephemeral Dataproc clusters. This matches the requirement for a visual interface with marketplace transforms.

Exam trap

The trap is confusing Dataprep with Data Fusion — both are visual, but Dataprep is for interactive data preparation while Data Fusion is the full ETL pipeline builder with a transform marketplace.

How to eliminate wrong answers

Option A is wrong because Dataproc is a managed Hadoop/Spark service that requires you to write code (Spark, Hive, Pig) — it provides infrastructure, not a drag-and-drop ETL designer. Option C is wrong because Dataprep (by Trifacta) is a visual data-wrangling tool for exploring and preparing data, but it is oriented toward interactive data preparation rather than full ETL pipeline orchestration with a transform marketplace. Option D is wrong because BigQuery is a serverless data warehouse for storing and querying data with SQL; it is a destination/analytics engine, not a visual ETL pipeline builder.

416
MCQmedium

A retail company processes real-time clickstream data using Cloud Pub/Sub and Dataflow. The pipeline aggregates events by user session and writes to Bigtable for low-latency queries. However, users report that session data is sometimes missing or duplicated. What is the most likely cause?

A.Session windowing is configured with too short a gap duration.
B.Bigtable schema design causes row key collisions.
C.Dataflow's default behavior discards late-arriving data.
D.Pub/Sub provides at-least-once delivery, and Dataflow does not deduplicate by default.
AnswerD

Pub/Sub guarantees at-least-once delivery, so messages can be redelivered after acknowledgement timeouts or retries. Dataflow's default pipeline performs no deduplication, so these redeliveries produce duplicate session records, while late or dropped acknowledgements also cause apparent gaps.

Why this answer

D is correct because Pub/Sub offers at-least-once delivery, meaning the same message may be delivered multiple times. Dataflow does not automatically deduplicate messages unless explicitly configured (e.g., using idempotent sinks or custom deduplication logic). Without deduplication, the same session event can be processed more than once, leading to duplicate session data in Bigtable.

Exam trap

Google Cloud often tests the misconception that Pub/Sub provides exactly-once delivery or that Dataflow automatically deduplicates messages from Pub/Sub, when in fact Pub/Sub is at-least-once and Dataflow requires explicit deduplication for idempotent processing.

How to eliminate wrong answers

Option A is wrong because a short gap duration would cause sessions to be split prematurely, leading to missing data (events not grouped into the same session), not duplicates. Option B is wrong because row key collisions in Bigtable would cause overwrites or errors, not missing or duplicate session data; Bigtable uses lexicographic ordering and row keys are unique per write. Option C is wrong because Dataflow's default behavior for late-arriving data depends on the windowing strategy; with session windows, late data can be included if within the allowed lateness, and Dataflow does not discard late data by default—it uses a default allowed lateness of 0 seconds, which would cause late data to be dropped, but this would result in missing data, not duplicates.

417
MCQeasy

A data engineer wants to store archived log files in Cloud Storage with a retention policy that prevents deletion for 5 years. Which feature should they use?

A.Object Lifecycle rule with Delete action after 5 years
B.Retention Policy on the bucket set to 5 years
C.Object Hold (temporal)
D.Versioning enabled
AnswerB

A bucket retention policy locks objects for a fixed period, preventing deletion or modification until it expires. Setting it to five years satisfies the stem's immutability requirement, unlike lifecycle rules, which delete or transition objects and therefore cannot enforce retention.

Why this answer

A retention policy on a Cloud Storage bucket enforces a minimum retention period for all objects in the bucket, preventing deletion or overwrite until the policy duration has elapsed. Setting it to 5 years ensures that archived log files cannot be deleted before that time, meeting the data engineer's requirement exactly. This is a bucket-level, immutable setting that applies to all objects, unlike object-level holds or lifecycle rules.

Exam trap

Google often tests the distinction between lifecycle rules that delete objects and retention policies that prevent deletion, so the trap here is assuming that a lifecycle rule with a Delete action can enforce a retention period, when in fact it does the opposite.

How to eliminate wrong answers

Option A is wrong because an Object Lifecycle rule with a Delete action after 5 years would automatically delete objects after 5 years, which is the opposite of preventing deletion; it does not enforce a retention period. Option C is wrong because an Object Hold (temporal) is a temporary hold placed on individual objects for a specific duration (e.g., days), not a bucket-wide policy for 5 years, and it is typically used for legal or compliance holds, not long-term retention. Option D is wrong because Versioning enabled preserves previous versions of objects but does not prevent deletion of the current version; it allows recovery after deletion but does not block deletion itself, so it does not enforce a retention policy.

418
MCQeasy

A financial services firm stores customer transaction data in a BigQuery table. The table contains a column `customer_id` that is frequently used in WHERE clauses, but the table is not partitioned. Queries filtering on a specific `customer_id` scan the entire table, which is large and costly. The data engineer wants to reduce the bytes scanned for these queries without changing the table's partitioning scheme. What should the engineer do?

A.Enable the `require_partition_filter` option on the table and add a partition on a date column.
B.Create a materialized view that selects all columns and filters on `customer_id`.
C.Add a partition on `customer_id` using a CREATE TABLE ... PARTITION BY customer_id statement.
D.Create a clustered table on `customer_id` by using a CREATE TABLE ... CLUSTER BY customer_id statement and loading the data into it.
AnswerD

Clustering sorts the data by the specified column and stores it in blocks, allowing BigQuery to prune unnecessary blocks when a query filters on that column. This reduces the bytes scanned and improves performance for queries that filter on `customer_id`. Since the table is not partitioned, clustering is the appropriate technique to achieve the goal without altering the partitioning scheme.

Why this answer

Clustering on `customer_id` organizes the table data so that BigQuery can skip blocks that do not contain the requested customer ID, thereby reducing bytes scanned and improving query performance. Partitioning is not suitable for a high-cardinality identifier like `customer_id`, and materialized views or partition filters do not address the specific need. Clustering is the correct technique for optimizing filters on a non-partitioned table.

Exam trap

The trap here is confusing clustering with partitioning, and assuming that any column can be used as a partition key, when in fact partitioning is limited to date/timestamp or integer range columns and high-cardinality columns like customer_id are best suited for clustering.

419
Multi-Selecteasy

Which TWO are benefits of using Vertex AI Endpoints for model serving?

Select 2 answers
A.Batch prediction support out of the box.
B.Integrated monitoring for prediction latency and error rates.
C.Automatic scaling based on traffic.
D.Automatic model retraining when drift is detected.
E.Built-in support for A/B testing without any additional configuration.
AnswersB, C

Vertex AI Endpoints provide built-in Cloud Monitoring metrics covering prediction latency, error rates and request counts per deployed model, satisfying the stem's requirement for serving benefits. This native observability removes the need to instrument custom telemetry, letting operators detect degradation and trigger autoscaling responses directly from the endpoint's own dashboards.

Why this answer

Option B is correct because Vertex AI Endpoints provide integrated monitoring that surfaces prediction latency and error rates, letting you track online serving health and performance directly in the console. Option C is correct because Vertex AI Endpoints support automatic scaling (autoscaling) based on traffic, adjusting the number of deployed model replicas to match request load. Options A, D, and E are not benefits of Endpoints: batch prediction is a separate Vertex AI feature (Batch Prediction jobs) rather than an Endpoint capability, automatic retraining on drift is not a built-in Endpoint function, and A/B testing requires explicit traffic-split configuration rather than working with no additional setup.

Exam trap

Google Cloud often tests the distinction between features that are 'built-in' versus those that require separate services or additional configuration, so candidates mistakenly assume batch prediction or automatic retraining are part of Endpoints when they are actually separate Vertex AI components.

420
MCQmedium

Your team uses Cloud Composer to orchestrate a daily ETL workflow that extracts data from an on-premises database, transforms it in Dataproc, and loads it into BigQuery. The workflow occasionally fails due to transient network issues. You want to automatically retry failed tasks without manual intervention. What should you do?

A.Configure the Dataproc cluster to automatically restart on failure.
B.Enable automatic retries on the Cloud Composer environment.
C.Use Cloud Scheduler to trigger the DAG multiple times until it succeeds.
D.Set the retries parameter and retry_delay on the Airflow tasks in the DAG.
AnswerD

In Cloud Composer, Airflow tasks support retries and retry_delay parameters. By setting these, you can automatically retry failed tasks after a specified delay. This is the standard way to handle transient failures in Airflow DAGs. You can also set exponential backoff for retries. This approach requires no additional infrastructure and is fully integrated with Cloud Composer.

Why this answer

Airflow tasks in Cloud Composer support retries and retry_delay parameters, which allow automatic retries of failed tasks. This is the correct way to handle transient failures. Other options either do not address task-level retries or could cause data duplication.

Exam trap

The trap here is thinking that Cloud Composer has an environment-wide retry setting, but retries must be configured per task.

421
MCQhard

You are preparing data in BigQuery for analysis. You have a table `orders` with columns `order_id`, `customer_id`, `order_date`, and `amount`. You need to create a new table that includes all orders, plus a column `prev_order_amount` that contains the amount of the customer's previous order by `order_date`. If there is no previous order, the value should be NULL. Which SQL feature should you use?

A.Use the LAG window function partitioned by `customer_id` and ordered by `order_date`.
B.Use the LEAD window function partitioned by `customer_id` and ordered by `order_date`.
C.Use a self-join on `customer_id` where the previous order date is the maximum order date less than the current order date.
D.Use the FIRST_VALUE window function partitioned by `customer_id` and ordered by `order_date`.
AnswerA

The LAG window function accesses data from a previous row in the same result set without the need for a self-join. Partitioning by `customer_id` ensures that the previous order is for the same customer, and ordering by `order_date` defines the sequence. LAG(amount) OVER (PARTITION BY customer_id ORDER BY order_date) returns the amount of the previous order, or NULL if there is no previous row. This is the most efficient and readable solution.

Why this answer

The LAG window function is designed to access a previous row's value within a partition. By partitioning by customer and ordering by order date, it returns the amount of the customer's previous order. This is efficient and avoids self-joins.

Other window functions like LEAD look forward, and FIRST_VALUE returns the first value in the partition, not the previous row.

Exam trap

The trap here is confusing LAG with LEAD, or thinking a self-join is necessary for previous-row access, when window functions are the intended solution.

422
Multi-Selectmedium

A team runs a Dataflow batch pipeline that writes to BigQuery. They want to reduce cost and improve reliability of the write step. The pipeline currently uses BigQueryIO with the default write method. Which two changes should they make to achieve these goals? (Choose two.)

Select 2 answers
A.Increase the number of BigQuery load jobs by setting a very low triggering frequency so each micro-batch is written separately.
B.Set the BigQueryIO write disposition to WRITE_TRUNCATE on every micro-batch so the table is always consistent.
C.Enable the withAutoSharding option on the write transform so the number of shards adapts to the pipeline's throughput.
D.Disable dynamic work rebalancing so that workers process a fixed set of shards for the entire pipeline run.
E.Switch the write to the Storage Write API with exactly-once semantics instead of the legacy streaming inserts or load jobs.
AnswersC, E

withAutoSharding dynamically adjusts the number of shards used for the BigQuery write based on the pipeline's throughput, which improves performance and reduces the chance of a single shard becoming a bottleneck. It is compatible with the Storage Write API and helps the write step scale without manual tuning. This supports reliability and cost efficiency by avoiding over-provisioning of shards for low-volume periods.

Why this answer

The Storage Write API lowers ingestion cost and supports exactly-once semantics, addressing both cost and reliability. Enabling withAutoSharding lets the write step adapt shard count to actual throughput, preventing shard bottlenecks and avoiding over-provisioning. Together these changes reduce per-row cost and improve the robustness of the BigQuery write step without altering the pipeline's business logic.

Exam trap

The trap here is assuming that more frequent, smaller load jobs improve reliability, when they actually increase quota pressure and metadata overhead while raising cost.

423
Multi-Selecthard

A team is deploying a complex model with multiple preprocessing steps. They want to ensure consistent preprocessing during training and serving. Which three approaches can achieve this? (Select 3)

Select 3 answers
A.Store preprocessing logic in a shared Python module
B.Use a separate preprocessing service called from the model
C.Use two separate pipelines for training and serving
D.Use Vertex AI Feature Transform Engine
E.Embed preprocessing logic in the model graph
AnswersA, D, E

A shared Python module guarantees identical preprocessing code executes during both training and serving, eliminating transformation drift. It directly satisfies the consistency constraint by centralising logic in one versioned artefact that both pipelines import, rather than duplicating steps. This is the standard mechanism for keeping feature engineering reproducible across environments.

Why this answer

Option A is correct because placing preprocessing logic in a shared Python module lets both the training code and the serving code import and execute the exact same transformation functions, eliminating drift between environments. Option D is correct because Vertex AI Feature Transform Engine lets you define transformations (e.g., in a feature store or as part of a Vertex AI pipeline) that are applied consistently at both training and online/batch serving time. Option E is correct because embedding preprocessing logic directly in the model graph (for example, as tf.keras preprocessing layers or a tf.transform graph) ensures the saved model performs the same transformations during inference as were applied during training.

Option B is not appropriate because a separate preprocessing service introduces an extra network hop, latency, and a potential point of failure, and it does not by itself guarantee identical logic between training and serving. Option C is incorrect because maintaining two separate pipelines for training and serving is precisely the practice that causes training-serving skew.

Exam trap

Google Cloud often tests the misconception that a separate preprocessing service (Option B) is a good architectural pattern for consistency, when in fact it introduces a single point of failure and versioning complexity that undermines the goal of identical preprocessing.

424
MCQmedium

A retail company runs a nightly batch pipeline that loads point-of-sale transactions into BigQuery. The pipeline uses a Cloud Composer (Apache Airflow) DAG with a BigQueryInsertJobOperator task that runs a SQL MERGE statement to upsert yesterday's sales into a large fact_sales table partitioned by transaction_date. The data engineering team notices that the MERGE task occasionally fails with a 'Resources exceeded during query execution' error when the source staging table contains more than 50 million rows. They need to redesign the DAG to reliably load large daily volumes without changing the final table schema or downstream dashboards. What should they do?

A.Replace the single MERGE statement with a multi-statement script that first deletes rows from the target partition for the affected dates and then inserts the staged rows using a BigQueryInsertJobOperator task with writeDisposition set to WRITE_APPEND.
B.Schedule the MERGE task to run during off-peak hours and increase the BigQuery slot reservation for the project so that the MERGE has more compute resources available.
C.Increase the number of Dataflow workers in the DAG by adding a DataflowOperator task that reads from the staging table and writes to the fact_sales table using a BigQueryIO write with STREAMING_INSERTS.
D.Convert the fact_sales table to a clustered table on transaction_date and re-run the original MERGE statement, relying on clustering to reduce the bytes scanned and avoid the resource error.
AnswerA

This approach avoids the memory-intensive shuffle of a single large MERGE by splitting work into a bounded DELETE on a partition and a streaming-friendly INSERT. BigQuery can execute each statement with less resource contention, and WRITE_APPEND avoids rewriting the entire target table. It preserves the schema and downstream behavior while scaling to large daily volumes.

Why this answer

A single large MERGE in BigQuery can hit internal resource limits because it must shuffle and join all source and target rows. Splitting the operation into a partition-scoped DELETE followed by an INSERT with WRITE_APPEND reduces the working set per statement and avoids the expensive shuffle. This pattern is a common, reliable way to upsert large daily batches while keeping the target schema and downstream queries unchanged.

Exam trap

The trap here is assuming that adding slots or clustering will fix a query that fails due to internal shuffle/memory limits, when the real fix is to decompose the MERGE into smaller, partition-scoped operations.

425
Multi-Selectmedium

You need to ingest streaming data from a custom application into BigQuery with exactly-once semantics and low latency. The data volume is up to 10 MB/s. Which TWO services should you combine?

Select 2 answers
A.Pub/Sub
B.Cloud Functions
C.BigQuery legacy streaming inserts
D.Dataflow with Storage Write API
E.Datastream
AnswersA, D

Pub/Sub provides the messaging layer that satisfies the exactly-once delivery constraint: its subscription-level exactly-once delivery feature deduplicates redelivered messages using acknowledgement IDs, so each record reaches BigQuery once. It also sustains the 10 MB/s throughput with low latency, unlike batch-oriented alternatives.

Why this answer

Option A (Pub/Sub) is correct because it provides a highly scalable, low-latency messaging service that decouples the custom application from downstream processing, ingesting up to 10 MB/s and beyond while buffering the stream reliably. Option D (Dataflow with Storage Write API) is correct because Dataflow can read from Pub/Sub and write to BigQuery using the Storage Write API, which supports exactly-once semantics via its exactly-once delivery mode and provides low-latency, high-throughput ingestion. Together, Pub/Sub plus Dataflow with the Storage Write API form the recommended architecture for exactly-once streaming ingestion into BigQuery.

Option B (Cloud Functions) is not appropriate here because it is event-driven and not designed for sustained high-throughput streaming pipelines with exactly-once guarantees. Option C (BigQuery legacy streaming inserts) does not provide exactly-once semantics and is being deprecated in favor of the Storage Write API. Option E (Datastream) is a change data capture (CDC) and replication service for databases, not a general-purpose streaming ingestion path for custom application data.

Exam trap

PDE often tests the misconception that BigQuery legacy streaming inserts provide exactly-once semantics, when in fact they only offer at-least-once delivery and can duplicate data; candidates must recognize that the Storage Write API is required for exactly-once.

426
Multi-Selectmedium

A data engineer is designing a BigQuery table for time-series data that will be queried frequently by time range and also by a customer_id. Which TWO design decisions will improve query performance and manage costs? (Choose two.)

Select 2 answers
A.Partition the table by day on the timestamp column
B.Cluster the table on customer_id
C.Disable automatic reclustering to save costs
D.Set partition expiration to 1 year
E.Use nested repeated fields for customer data
AnswersA, B

Enables partition pruning for time-range queries.

Why this answer

Partitioning the table by day on the timestamp column allows BigQuery to prune partitions when queries filter by a time range, scanning only the relevant partitions instead of the entire table. This directly reduces the amount of data read, improving query performance and lowering costs.

Exam trap

Google Cloud often tests the misconception that disabling automatic reclustering saves costs, but in reality it is free and essential for maintaining clustering benefits, while partition expiration is a lifecycle management feature, not a performance optimization.

427
MCQmedium

A company stores sensitive data in Cloud Storage and must ensure that data is encrypted at rest with keys that they control and can rotate on demand. They also need to audit key usage and revoke access immediately if a key is compromised. Which Cloud Storage encryption option should they use?

A.Google-managed encryption keys
B.Customer-supplied encryption keys (CSEK)
C.Client-side encryption before uploading to Cloud Storage
D.Customer-managed encryption keys (CMEK) with Cloud KMS
AnswerD

CMEK with Cloud KMS lets the company create, rotate, and manage keys in Cloud KMS, and Cloud Storage uses these keys to encrypt data at rest. Cloud KMS provides audit logs for key usage and allows immediate revocation by disabling or destroying the key, meeting all requirements for control, audit, and revocation.

Why this answer

Customer-managed encryption keys (CMEK) with Cloud KMS allow the company to control and rotate keys, audit key usage through Cloud KMS audit logs, and revoke access by disabling or destroying the key. This satisfies the requirements for customer-controlled encryption, on-demand rotation, and immediate revocation for sensitive Cloud Storage data.

Exam trap

The trap here is assuming that customer-supplied encryption keys give the same integrated key management, rotation, and audit capabilities as CMEK, when CSEK keys are not stored or managed by Google Cloud.

428
MCQeasy

Which Google Cloud database offers global distribution, strong consistency, and a 99.999% SLA?

A.Cloud Spanner
B.Cloud Bigtable
C.Firestore
D.Cloud SQL
AnswerA

Cloud Spanner synchronously replicates across regions using TrueTime, delivering externally consistent reads globally plus a 99.999% availability SLA. Regional Cloud SQL and Firestore cannot match that combination, and Bigtable offers neither strong consistency nor that SLA.

Why this answer

Cloud Spanner is the only Google Cloud database that provides global distribution (horizontally scaling across regions), strong consistency (external consistency with TrueTime), and a 99.999% SLA. It combines the benefits of relational database structure with non-relational horizontal scale, making it ideal for globally distributed, strongly consistent workloads.

Exam trap

Candidates may confuse Firestore's multi-region eventual consistency with the strong consistency required for the 99.999% SLA, or assume Cloud Bigtable's high throughput implies strong consistency. Spanner's external consistency with TrueTime is unique to Google Cloud.

How to eliminate wrong answers

Option B (Cloud Bigtable) is wrong because it offers only eventual consistency (not strong consistency) and a 99.99% SLA, not 99.999%. Option C (Firestore) is wrong because it provides strong consistency only within a single region; its multi-region mode uses eventual consistency, and its SLA is 99.999% only for single-region, not globally distributed strong consistency. Option D (Cloud SQL) is wrong because it is a single-region relational database with no global distribution capability and a 99.95% SLA.

429
MCQhard

A company has a batch prediction job that runs daily using AI Platform Batch Prediction. The job uses a TensorFlow model and processes 10 GB of data. Recently, the job started failing with the error 'The replica worker 0 exited with a non-zero exit code: Out of memory'. Which action should the team take to resolve this without rewriting the model?

A.Increase the number of workers (parallelism) to distribute the data across more machines.
B.Use a machine type with more memory, such as n1-highmem-8.
C.Reduce the batch size parameter in the prediction job configuration.
D.Optimize the model to use less memory by pruning or quantization.
AnswerB

The out-of-memory error originates from the replica worker exhausting RAM during batch prediction, not from the model itself. Selecting a higher-memory machine type such as n1-highmem-8 gives the worker sufficient memory, resolving the failure without altering the TensorFlow model.

Why this answer

The error 'Out of memory' on replica worker 0 indicates that the machine type assigned to the prediction job does not have enough RAM to load the model and process the 10 GB batch. Increasing the machine type to one with more memory (e.g., n1-highmem-8) directly addresses the memory constraint without requiring any code changes. This is the most straightforward fix because AI Platform Batch Prediction allows you to specify machine types in the job configuration, and the error is purely a resource allocation issue.

Exam trap

Google Cloud often tests the distinction between scaling horizontally (adding workers) and scaling vertically (increasing machine resources), where candidates mistakenly assume parallelism solves memory issues, but the error is per-worker memory exhaustion, not throughput.

How to eliminate wrong answers

Option A is wrong because increasing the number of workers (parallelism) distributes the data across more machines but does not increase the memory per worker; each replica still has the same limited memory, so the out-of-memory error would persist on each worker. Option C is wrong because reducing the batch size parameter controls how many predictions are processed per step, which can reduce peak memory usage per request, but the error occurs during model loading or initial data processing, not during per-step prediction; the 10 GB dataset and model size still require sufficient base memory. Option D is wrong because while pruning or quantization could reduce model memory footprint, the question explicitly states 'without rewriting the model,' and these techniques require modifying the model architecture or retraining, which is a form of rewriting.

430
MCQmedium

A team is designing a data lake on Google Cloud using Cloud Storage and BigQuery. They need to ensure that sensitive data (e.g., PII) is encrypted at rest and have the ability to audit access. Which approach meets these requirements?

A.Use Customer-Managed Encryption Keys (CMEK) and enable VPC Service Controls.
B.Use Customer-Managed Encryption Keys (CMEK) and enable Cloud Audit Logs.
C.Use Default Encryption and enable Data Loss Prevention (DLP) API.
D.Use Customer-Supplied Encryption Keys (CSEK) and enable VPC Service Controls.
AnswerB

CMEK lets the team control and rotate the encryption keys protecting PII at rest, while Cloud Audit Logs record every access to that data. Together they satisfy both the encryption-at-rest and access-auditing requirements for the Cloud Storage and BigQuery data lake.

Why this answer

Customer-Managed Encryption Keys (CMEK) allow the team to control and manage the encryption keys used to protect data at rest in Cloud Storage and BigQuery, while enabling Cloud Audit Logs provides the necessary audit trail for access to both the data and the keys. This combination directly satisfies the requirements for encryption at rest and auditability.

Exam trap

Google Cloud often tests the distinction between encryption key management (CMEK vs. CSEK vs. Default) and security controls (VPC Service Controls vs.

Audit Logs), leading candidates to conflate network perimeter controls with audit capabilities.

How to eliminate wrong answers

Option A is wrong because VPC Service Controls provide network-based security boundaries to prevent data exfiltration, but they do not provide audit logging of access to data or keys, which is a separate requirement. Option C is wrong because Default Encryption uses Google-managed keys, which do not give the team control over encryption keys, and the DLP API is for inspecting and classifying sensitive data, not for encryption at rest or audit logging. Option D is wrong because Customer-Supplied Encryption Keys (CSEK) require the customer to manage their own keys outside Google Cloud, which adds operational complexity and does not integrate with Cloud Audit Logs for key access auditing; VPC Service Controls again do not provide audit logging.

431
MCQeasy

A team needs to orchestrate a multi-step workflow that involves calling external APIs, running BigQuery queries, and conditionally executing Cloud Functions. Which Google Cloud service is best suited for this?

A.Dataflow
B.Workflows
C.Cloud Composer
D.Cloud Scheduler
AnswerB

Workflows orchestrates multi-step processes using YAML or JSON definitions, sequencing HTTP calls to external APIs, BigQuery jobs, and conditional Cloud Functions invocations. It directly satisfies the stem's requirement for conditional branching across heterogeneous services, unlike single-purpose tools such as Cloud Scheduler or Pub/Sub, which cannot express stateful multi-step logic.

Why this answer

Workflows is a serverless orchestration service that allows you to define multi-step workflows as a sequence of steps, including HTTP calls to external APIs, BigQuery queries, and conditional logic to invoke Cloud Functions. It integrates natively with other Google Cloud services via the Workflows API and supports error handling, retries, and parallel steps, making it ideal for this use case.

Exam trap

Candidates often confuse orchestration services (Workflows) with data processing services (Dataflow) or scheduling services (Cloud Scheduler), leading them to choose Dataflow because they mistake data processing for workflow orchestration.

How to eliminate wrong answers

Option A is wrong because Dataflow is a stream and batch data processing service based on Apache Beam, not an orchestration tool for coordinating API calls, BigQuery queries, and Cloud Functions. Option C is wrong because Cloud Composer is a managed Apache Airflow service that is designed for complex, scheduled workflows with dependencies, but it is heavier, requires more setup, and is overkill for a simple multi-step orchestration that Workflows handles more efficiently. Option D is wrong because Cloud Scheduler is a cron job service for triggering tasks on a schedule, but it cannot orchestrate conditional logic, API calls, or BigQuery queries within a single workflow.

432
MCQhard

You are migrating a large on-premises data warehouse to BigQuery. The data includes sensitive PII columns that must be masked for certain users. Which BigQuery feature can automatically redact PII in query results based on user roles?

A.IAM conditions on tables
B.Authorized views
C.Column-level security with data masking
D.Cloud DLP API
AnswerC

Column-level security with data masking attaches policy tags to sensitive PII columns, so BigQuery automatically redacts or hashes those values in query results for users lacking the required role, satisfying the masking requirement without rewriting queries.

Why this answer

BigQuery column-level security with data masking (also called dynamic data masking) lets you attach policy tags to specific columns and define masking rules that automatically redact values in query results based on the querying user's IAM roles. Users without the fine-grained reader role see masked values (e.g., hashed or null), while authorized users see the real data — all without changing the query. This is the native BigQuery feature designed for exactly this PII-redaction use case.

Exam trap

The trap is confusing Cloud DLP (a data discovery and de-identification service) with BigQuery's native column-level security and dynamic data masking, which is the feature that redacts query results based on user roles at query time.

How to eliminate wrong answers

Option A is wrong because IAM conditions on tables control who can access the table at all, not how individual columns are masked in results — they are coarse-grained access control. Option B is wrong because authorized views require you to create and maintain separate views for each access pattern and do not automatically redact PII based on roles. Option D is wrong because Cloud DLP API is a separate service for discovering, classifying, and de-identifying sensitive data (e.g., during ETL or via DLP jobs), but it does not automatically redact query results based on user roles at query time.

433
MCQhard

A financial services company needs to process credit card transactions in real time to detect fraudulent patterns. The pipeline must handle late-arriving data (up to 2 hours) and produce accurate results. They want to use a unified programming model that works for both batch and streaming. Which Google Cloud service should they use?

A.Cloud Pub/Sub with Cloud Functions
B.BigQuery with scheduled queries
C.Cloud Dataproc with Spark Streaming
D.Cloud Dataflow with Apache Beam
AnswerD

Cloud Dataflow, based on Apache Beam, provides a unified programming model for batch and streaming. It supports event-time processing, watermarks, and triggers to handle late-arriving data accurately. With windowing and allowed lateness, it can produce correct results even when data arrives up to 2 hours late, making it ideal for real-time fraud detection with late data.

Why this answer

Cloud Dataflow with Apache Beam is the correct choice because it offers a unified batch and streaming programming model, supports event-time processing with watermarks and triggers, and can handle late-arriving data accurately through configurable allowed lateness. This makes it well-suited for real-time fraud detection with data arriving up to 2 hours late.

Exam trap

The trap here is assuming that any streaming service can handle late data, but only Dataflow with Beam provides built-in event-time semantics and allowed lateness for accurate results.

434
MCQmedium

You configured a model deployment monitor on your Vertex AI endpoint as shown. What will happen when the feature 'age' has a skew of 0.4?

A.An alert will be sent to admin@example.com
B.The endpoint will automatically roll back to a previous model version
C.No alert will be sent because the skew threshold is 0.2 for income
D.An alert will be sent only if both features exceed their thresholds
AnswerA

The deployment monitor's configured skew threshold is exceeded by the 0.4 value, triggering the notification action bound to that monitor. Because admin@example.com is the registered alert recipient, the breach fires an email alert to that address.

Why this answer

The monitoring configuration shows an alert threshold of 0.3 for the feature 'age', and a skew of 0.4 exceeds that threshold. Vertex AI Model Monitoring will trigger the configured alert action, which in this case is sending an email to admin@example.com. The alert is based on the specific feature's threshold, not on any other feature's threshold.

Exam trap

Google Cloud often tests the misconception that alerts require multiple features to exceed thresholds or that the system can automatically roll back models, when in reality each feature is evaluated independently and only notifications are sent.

How to eliminate wrong answers

Option B is wrong because Vertex AI Model Monitoring does not automatically roll back model deployments; it only sends alerts based on configured actions, and auto-rollback is not a supported feature in this context. Option C is wrong because the skew threshold for 'age' is 0.2, not 0.2 for 'income'; the question states the skew for 'age' is 0.4, which exceeds its own threshold, so an alert will be sent regardless of the 'income' feature's threshold. Option D is wrong because the alert is triggered per feature when its individual threshold is exceeded; there is no requirement for both features to exceed their thresholds simultaneously.

435
MCQhard

A financial firm ingests trade confirmations from a partner's SFTP server into Cloud Storage, then loads them into BigQuery. The partner drops a variable number of files at unpredictable times, and the firm must detect each new file within seconds and trigger a downstream Cloud Run job. The files must not be processed twice if the trigger fires more than once. Which design should you implement?

A.Schedule a Cloud Scheduler job every minute that lists the bucket and invokes Cloud Run for any object newer than the last run.
B.Use a Pub/Sub topic populated by Cloud Storage notifications and subscribe with a pull subscription that Cloud Run polls on a schedule.
C.Attach a Cloud Functions (1st gen) trigger to the bucket and let it call Cloud Run synchronously for each new object.
D.Configure an Eventarc trigger on the bucket's object.finalize events to invoke Cloud Run, and have the function record the object generation in Firestore before processing.
AnswerD

Eventarc delivers Cloud Storage object.finalize events within seconds of upload, satisfying the latency requirement. Because at-least-once delivery can produce duplicate events, recording the object's generation number as an idempotency key before doing work ensures a repeated trigger for the same object version is skipped, preventing double processing.

Why this answer

Eventarc on object.finalize gives sub-minute, event-driven detection without polling, and the object generation number serves as a natural idempotency key because it uniquely identifies a version of the object. Writing that key to Firestore before doing the work makes duplicate deliveries harmless. Polling designs miss the latency target, and designs without a deduplication record allow the same file to be processed again on redelivery.

Exam trap

The trap here is assuming that a Cloud Storage event fires exactly once per object, when delivery is at-least-once and the same object can generate repeated triggers that must be reconciled by a durable idempotency key.

436
Drag & Dropmedium

Drag and drop the steps to migrate an on-premises MySQL database to Cloud SQL using Database Migration Service into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Database Migration Service enables minimal-downtime migrations using replication.

437
Multi-Selecthard

A company uses Cloud Composer to orchestrate Dataproc and BigQuery jobs. They need to implement retry logic for transient failures. Which THREE features can help?

Select 3 answers
A.Dataflow pipeline retries
B.DAG retry_delay
C.BigQuery job retries
D.Cloud Composer high availability
E.Task retries and retry_delay
AnswersB, C, E

Composer can retry the entire DAG on failure with a delay.

Why this answer

Cloud Composer (Apache Airflow) allows setting `retry_delay` at the DAG level to define the time delay between task retries. This is a native Airflow feature that helps handle transient failures by automatically retrying failed tasks after a specified delay, reducing manual intervention.

Exam trap

The trap here is confusing infrastructure-level high availability (Option D) with application-level retry logic, leading candidates to select HA as a retry mechanism when it only ensures environment uptime, not task-level failure recovery.

438
MCQmedium

An e-commerce company runs a daily batch pipeline that processes clickstream data from Cloud Storage using Cloud Dataproc with Spark. The pipeline includes a join between a large fact table and a small dimension table. The dimension table is stored in Cloud Storage as a CSV file. The join is slow due to shuffling. The data engineer considers broadcasting the dimension table. However, the dimension table is updated daily and the pipeline reads the latest version. What is the best approach to implement this optimization?

A.Use DataFrame.join with broadcast hint on the dimension DataFrame
B.Read the fact table and dimension table into separate DataFrames and use standard join
C.Read the dimension table as an RDD and collect as a map, then use map-side join
D.Increase the spark.sql.autoBroadcastJoinThreshold to a large value
AnswerA

Broadcast hints avoid the shuffle by replicating the small dimension table to each executor, eliminating the expensive join shuffle. Reading the CSV fresh each run keeps the daily-updated dimension current, satisfying the freshness constraint without caching stale data.

Why this answer

Broadcasting the small dimension table using the broadcast hint (e.g., `broadcast(dimensionDF)`) forces Spark to replicate the dimension data to all executor nodes, eliminating the need for a shuffle during the join. This is ideal when the dimension table is small enough to fit in executor memory, and since the pipeline reads the latest CSV daily, the broadcast will automatically use the updated data without additional code changes.

Exam trap

The trap here is that candidates may think increasing `spark.sql.autoBroadcastJoinThreshold` is a safe global fix, but it can cause memory pressure and does not guarantee a broadcast join if the table size fluctuates, whereas the explicit broadcast hint provides deterministic behavior.

How to eliminate wrong answers

Option B is wrong because a standard join without any hint or optimization will trigger a full shuffle of both datasets, which is exactly the performance problem described. Option C is wrong because manually collecting the dimension table as an RDD and using a map-side join is an outdated, error-prone approach that bypasses Spark SQL's Catalyst optimizer and broadcast join optimizations; it also requires manual handling of updates and memory management. Option D is wrong because increasing `spark.sql.autoBroadcastJoinThreshold` globally may cause the dimension table to be broadcast automatically, but it does not guarantee the join uses a broadcast if the table size exceeds the threshold, and it can lead to out-of-memory errors if the threshold is set too high without considering executor memory limits.

439
MCQeasy

A company uses Cloud Monitoring to track application latency. They notice a spike in latency every 30 minutes. What is the best initial step to diagnose the issue?

A.Increase the number of instances to handle the load.
B.Enable Cloud Trace for all requests.
C.Check if scheduled jobs or cron tasks overlap.
D.Change the alert threshold to ignore the spikes.
AnswerC

A latency spike recurring at a fixed 30-minute interval strongly suggests a periodic workload competing for resources. Checking scheduled jobs or cron tasks for overlap with the spike window is the fastest way to confirm or rule out that correlation before deeper analysis.

Why this answer

Recurring spikes at regular intervals often indicate a scheduled process (e.g., cron job, batch job) that runs every 30 minutes. Checking for overlapping scheduled jobs is the most efficient first step before scaling or other actions.

440
MCQmedium

A data engineer needs to store raw sensor data in Cloud Storage and automatically transition it to a lower-cost storage class after 30 days, then delete it after 365 days. What should they configure?

A.Use Cloud Pub/Sub notifications to trigger a Cloud Function that moves objects.
B.Use gsutil rewrite command in a cron job.
C.Configure a lifecycle rule with SetStorageClass to Nearline after 30 days and Delete after 365 days.
D.Set a bucket retention policy with a retention period of 365 days.
AnswerC

A single lifecycle rule with age-based conditions performs both transitions automatically: SetStorageClass to Nearline at 30 days, then Delete at 365 days. This satisfies both the cost-tiering and retention constraints without manual intervention or separate rules.

Why this answer

Cloud Storage lifecycle management rules allow you to automatically transition objects to a lower-cost storage class (such as Nearline) after a specified number of days and then delete them after another period. This is the native, serverless way to manage object lifecycle without external scripts or compute resources.

Exam trap

Google Cloud often tests the distinction between lifecycle management (which automates transitions and deletions) and retention policies (which only prevent deletion/overwrites), leading candidates to confuse the two.

How to eliminate wrong answers

Option A is wrong because Cloud Pub/Sub notifications and Cloud Functions introduce unnecessary complexity and cost; lifecycle rules handle this natively without custom code. Option B is wrong because using gsutil rewrite in a cron job is a manual, error-prone approach that does not scale and incurs additional egress/operation costs; lifecycle rules are the intended automated solution. Option D is wrong because a bucket retention policy prevents deletion before the retention period ends, but it does not automatically transition objects to a lower-cost storage class; it only enforces immutability.

441
MCQeasy

You need to allow a data analyst to run queries on a BigQuery dataset but prevent them from modifying the data or deleting the dataset. Which IAM role should you grant?

A.roles/bigquery.dataOwner
B.roles/bigquery.dataViewer
C.roles/bigquery.jobUser
D.roles/bigquery.dataEditor
AnswerB

roles/bigquery.dataViewer grants read access to datasets, tables and views, permitting queries while denying write operations such as inserts, updates, deletes or dataset removal. It therefore satisfies both constraints: the analyst can run queries but cannot modify data or delete the dataset.

Why this answer

roles/bigquery.dataViewer grants read-only access to BigQuery datasets, tables, and views, allowing the analyst to run queries and view metadata but not modify data or delete the dataset. It is the least-privilege role that satisfies the requirement of querying without write or delete permissions. This role is typically paired with roles/bigquery.jobUser at the project level so the user can actually run jobs.

Exam trap

The trap is picking roles/bigquery.jobUser alone because it 'lets you run queries,' but jobUser does not grant data access — you need dataViewer (or higher) on the dataset as well.

How to eliminate wrong answers

Option A is wrong because roles/bigquery.dataOwner grants full control over the dataset, including the ability to modify data, delete tables, and even delete the dataset — far more than the read-only access required. Option C is wrong because roles/bigquery.jobUser only allows the user to run jobs (queries, loads, exports) in a project; it does not by itself grant access to read the data in a dataset, so the analyst could not query the data. Option D is wrong because roles/bigquery.dataEditor allows the user to modify and delete data and tables, which violates the requirement to prevent modifications.

442
Multi-Selectmedium

Which THREE Google Cloud services are typically used together in a production ML pipeline?

Select 3 answers
A.Cloud Storage
B.Cloud Functions
C.Vertex AI Training
D.Vertex AI Prediction
E.BigQuery
AnswersA, C, D

For storing training data, model artifacts, etc.

Why this answer

Cloud Storage is correct because it serves as the central artifact repository in a production ML pipeline on Google Cloud. It stores training data, model artifacts, and prediction inputs/outputs, enabling seamless integration with Vertex AI Training for model training and Vertex AI Prediction for serving. Without Cloud Storage, there is no durable, scalable, and cost-effective way to manage the large datasets and model binaries required for production ML workflows.

Exam trap

The trap here is that candidates confuse 'services used in an ML pipeline' with 'services that can be used somewhere in ML' — Cloud Functions and BigQuery are often used in ML workflows (e.g., triggering retraining or storing features), but they are not the three core services that are typically used together in a production ML pipeline for training, storing artifacts, and serving predictions.

443
MCQmedium

A Dataflow pipeline reads log files from Cloud Storage, parses them into LogEvent objects, and writes to BigQuery. The pipeline fails with the above errors. What is the most likely cause?

A.The LogEvent class does not have a no-argument constructor.
B.The pipeline is missing required import statements for LogEvent.
C.The BigQuery table schema does not match the LogEvent fields.
D.The log files are not in the expected format, causing parsing failures.
AnswerA

Apache Beam serialises user-defined types between workers, and its default coder for a custom class requires a public no-argument constructor to instantiate the object during deserialisation. Without one, the pipeline cannot reconstruct LogEvent instances, producing the construction failure observed.

Why this answer

Apache Beam's SDK requires that custom types used as PCollection elements (like LogEvent) have a no-argument constructor so that the framework can deserialize objects during distributed processing, especially when using the Dataflow runner. Without it, the pipeline fails at runtime with a serialization error because Beam's default coder (e.g., SerializableCoder) cannot reconstruct the object.

Exam trap

The trap here is that candidates confuse runtime serialization errors with compile-time import issues or schema mismatches, overlooking the fundamental requirement for a no-argument constructor in Beam's default coders.

How to eliminate wrong answers

Option B is wrong because missing import statements would cause a compile-time error, not a runtime pipeline failure with the described errors. Option C is wrong because a BigQuery table schema mismatch would produce a write-time error (e.g., schema mismatch), not a serialization failure during parsing. Option D is wrong because parsing failures from malformed log files would result in exceptions during the parse step, not a serialization error related to the LogEvent class itself.

444
MCQhard

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle late-arriving data (up to 1 hour) and group events into 10-minute windows. Which configuration is correct?

A.Use global windows with a trigger that fires every 10 minutes
B.Use sliding windows of 10 minutes with a 5-minute period and allowed lateness of 1 hour
C.Use fixed windows of 10 minutes with allowed lateness of 0 seconds
D.Use fixed windows of 10 minutes with allowed lateness of 1 hour and a trigger that fires after watermark plus early firings
AnswerD

Fixed 10-minute windows satisfy the grouping requirement, while one hour of allowed lateness lets late events still update their window's results. The trigger firing on watermark plus early firings emits speculative results promptly, then corrects them once the watermark passes, so no late data within the hour is dropped.

Why this answer

Fixed windows of 10 minutes match the required grouping interval, allowed lateness of 1 hour accommodates late-arriving data up to one hour, and a trigger that fires after the watermark plus early firings ensures results are emitted promptly while still accepting late data. This combination is the canonical Apache Beam pattern for windowed aggregation with late data on a streaming pipeline.

Exam trap

PDE often tests whether candidates conflate windowing (grouping) with triggering (emission timing) — the trap is choosing a trigger-only answer (global windows with a periodic trigger) when the requirement is per-window aggregation with late data.

How to eliminate wrong answers

Option A is wrong because global windows do not partition events into 10-minute groups — a trigger firing every 10 minutes emits cumulative results over the entire stream, not per-window aggregates. Option B is wrong because sliding windows of 10 minutes with a 5-minute period produce overlapping windows, which is not what 'group events into 10-minute windows' requires and would duplicate aggregates. Option C is wrong because allowed lateness of 0 seconds discards all late data, directly violating the 1-hour late-arrival requirement.

445
Drag & Dropmedium

Drag and drop the steps to create a Cloud Storage bucket with uniform bucket-level access into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Uniform bucket-level access simplifies permissions by using IAM policies at the bucket level instead of ACLs.

446
MCQmedium

A company wants to use Cloud Data Fusion for ETL pipelines. They need to integrate with custom transformations not available in the marketplace. What should they do?

A.Switch to Dataproc and write a Spark job.
B.Use the Data Fusion Hub to download a custom plugin.
C.Use Dataprep to create the transformation.
D.Write a custom plugin using the CDAP SDK and deploy it.
AnswerD

The CDAP SDK lets developers author custom transformations as plugins, package them, and deploy into Cloud Data Fusion, extending the pipeline beyond marketplace offerings. This satisfies the requirement for transformations unavailable in the marketplace, which built-in operators and existing plugins cannot supply.

Why this answer

Cloud Data Fusion allows extending its capabilities by writing custom plugins using the CDAP SDK, which can then be deployed to the Data Fusion instance. This enables integration of transformations not available in the marketplace, providing full flexibility for custom ETL logic.

Exam trap

PDE often tests the extensibility of Cloud Data Fusion, and candidates might incorrectly assume that the Data Fusion Hub provides custom plugins or that other GCP services like Dataprep can be used interchangeably.

How to eliminate wrong answers

Option A is wrong because switching to Dataproc and writing a Spark job would abandon Cloud Data Fusion entirely, which is not necessary and would require re-architecting the pipeline. Option B is wrong because the Data Fusion Hub is a marketplace for pre-built plugins; it does not offer custom plugin downloads. Option C is wrong because Dataprep is a separate data preparation tool that does not integrate custom transformations into Data Fusion pipelines.

447
MCQeasy

You need to ingest Google Ads performance data into BigQuery on a daily basis for reporting. Which service should you use?

A.BigQuery Data Transfer Service for Google Ads
B.Cloud Scheduler to call Google Ads API and load to BigQuery
C.Pub/Sub with a Google Ads subscriber
D.Storage Transfer Service for Google Ads
AnswerA

BigQuery Data Transfer Service provides a managed, scheduled connector that pulls Google Ads performance data directly into BigQuery daily without custom extraction code. This satisfies the stem's daily ingestion requirement, handling authentication, incremental transfers and backfill automatically.

Why this answer

The BigQuery Data Transfer Service for Google Ads is the correct choice because it provides a fully managed, scheduled connector that automatically ingests Google Ads performance data into BigQuery on a daily basis without requiring any custom code. It handles authentication, schema mapping, and incremental loads, making it the simplest and most reliable solution for this specific use case.

Exam trap

Google often tests the distinction between fully managed services (like BigQuery Data Transfer Service) and generic infrastructure components (like Cloud Scheduler or Pub/Sub) that require custom development, leading candidates to overcomplicate the solution by choosing a more flexible but less appropriate option.

How to eliminate wrong answers

Option B is wrong because Cloud Scheduler is a cron job service that can trigger HTTP requests, but it does not natively integrate with the Google Ads API or handle the complex authentication, pagination, and schema mapping required to load data into BigQuery; you would still need to build and maintain a custom application. Option C is wrong because Pub/Sub is a messaging service for asynchronous event streaming, not a batch ingestion tool; while you could theoretically publish Google Ads data to Pub/Sub, there is no native Google Ads subscriber, and you would need to build a custom subscriber to write to BigQuery, which is far more complex than using the dedicated transfer service. Option D is wrong because Storage Transfer Service is designed for moving data from on-premises or cloud storage (like S3 or HTTP endpoints) into Google Cloud Storage, not for directly ingesting data from Google Ads into BigQuery.

448
Multi-Selecthard

A company is designing a data lake on Cloud Storage for analytics. They need to store data in various formats (Avro, Parquet, CSV) and enable efficient querying with BigQuery and Dataproc. Which THREE practices should they follow?

Select 3 answers
A.Use BigLake to create BigQuery tables that reference Cloud Storage data.
B.Store data in columnar formats like Parquet for analytics workloads.
C.Disable encryption on the bucket to improve read performance.
D.Partition data by date in a logical folder structure (e.g., /data/yyyy/mm/dd).
E.Store all data in CSV format for simplicity.
AnswersA, B, D

Enables querying data without loading.

Why this answer

BigLake allows you to create BigQuery tables that reference data stored in Cloud Storage, enabling unified governance and fine-grained access control without moving data. This is essential for a data lake architecture where BigQuery and Dataproc need to query the same underlying data in various formats like Avro, Parquet, and CSV.

Exam trap

Google Cloud often tests the misconception that disabling encryption improves performance, but Cloud Storage encryption is transparent and has no measurable impact on read throughput, so candidates should recognize that security controls are non-negotiable in cloud data lakes.

449
MCQeasy

A company uses Cloud Functions to process events from Cloud Storage. They notice that occasionally functions are not triggered. What should they check first to ensure solution quality?

A.Verify that the Cloud Storage bucket has notifications configured for the correct event type.
B.Check the logs for function execution.
C.Increase the function memory allocation.
D.Increase the function timeout.
AnswerA

Cloud Functions triggers on Cloud Storage depend on bucket notifications being configured for the correct event type (for example, object.finalize). Missing or mismatched notifications explain intermittent non-triggering, so verifying this configuration is the first diagnostic step.

Why this answer

The most common reason for Cloud Functions not being triggered by Cloud Storage events is that the bucket's notification configuration is missing or misconfigured. Cloud Functions relies on Pub/Sub notifications from the bucket to invoke the function; if the notification is not set for the correct event type (e.g., `OBJECT_FINALIZE`), the function will never be triggered. Therefore, verifying the notification configuration is the first and most direct diagnostic step.

Exam trap

Google often tests the misconception that performance tuning (memory or timeout) is the first step to fix trigger issues, when the root cause is almost always a missing or misconfigured event notification.

How to eliminate wrong answers

Option B is wrong because checking logs for function execution assumes the function was invoked, but if the trigger is not firing, there will be no execution logs to review. Option C is wrong because increasing memory allocation addresses performance issues like out-of-memory errors, not trigger failures. Option D is wrong because increasing the timeout addresses function execution duration limits, not the absence of invocation.

450
MCQeasy

An organization wants to automate their batch data processing pipeline using Cloud Composer. The pipeline consists of multiple tasks: extract from Cloud Storage, transform with Dataflow, and load into BigQuery. Which Airflow operator should be used to run Dataflow jobs?

A.BigQueryInsertJobOperator
B.DataflowCreatePythonJobOperator
C.GCSToBigQueryOperator
D.DataprocSubmitJobOperator
AnswerB

DataflowCreatePythonJobOperator launches a Python-defined Dataflow pipeline directly from Airflow, satisfying the requirement to run Dataflow jobs within the Composer DAG. It handles pipeline submission and job monitoring natively, unlike generic operators, making it the precise fit for the transform stage between Cloud Storage extraction and BigQuery loading.

Why this answer

B is correct because the DataflowCreatePythonJobOperator is specifically designed to submit and manage Apache Beam pipelines written in Python as Dataflow jobs in Google Cloud. This operator handles the creation of a Dataflow job from a Python file, which aligns with the requirement to run Dataflow transformations within a Cloud Composer DAG.

Exam trap

Google Cloud often tests the distinction between Dataflow and Dataproc operators, so the trap here is that candidates might confuse DataprocSubmitJobOperator (for Hadoop/Spark) with Dataflow operators, especially when the question mentions 'transform' without specifying the processing framework.

How to eliminate wrong answers

Option A is wrong because BigQueryInsertJobOperator is used to run BigQuery jobs (e.g., queries, load jobs), not to submit Dataflow pipelines. Option C is wrong because GCSToBigQueryOperator loads data directly from Cloud Storage to BigQuery without using Dataflow for transformation, bypassing the required transform step. Option D is wrong because DataprocSubmitJobOperator submits jobs to Dataproc (Hadoop/Spark clusters), not to Dataflow, which is a different processing service.

Page 5

Page 6 of 10

Page 7

All pages