Courseiva

CCNA Ingesting and Processing the Data Questions

32 of 107 questions · Page 2/2 · Ingesting and Processing the Data · Answers revealed

76
MCQmedium

A company uses Google Ads and wants to automatically load their advertising data into BigQuery daily. They also need to transform the data with SQL and schedule a recurring query. Which combination of services meets these requirements with minimal operational overhead?

A.Cloud Functions triggered by Cloud Scheduler to call Google Ads API and load into BigQuery
B.Cloud Composer to extract Google Ads API and Dataflow to transform
C.Storage Transfer Service to move CSV files to GCS, then load into BigQuery
D.BigQuery Data Transfer Service for Google Ads and scheduled queries
AnswerD

BigQuery Data Transfer Service natively ingests Google Ads data on a daily schedule, and BigQuery scheduled queries run the SQL transformations recurringly. Together they satisfy the daily load, SQL transformation, and scheduling requirements with no servers or orchestration code to maintain.

Why this answer

BigQuery Data Transfer Service has a built-in Google Ads connector that automatically loads advertising data into BigQuery on a schedule, and BigQuery scheduled queries let you transform that data with SQL on a recurring basis — both fully managed with no infrastructure. This combination meets the daily load and SQL transformation requirements with minimal operational overhead.

Exam trap

PDE often tests whether candidates choose the fully managed native connector (BigQuery DTS) over custom code (Cloud Functions/Composer) when the question emphasizes 'minimal operational overhead.'

How to eliminate wrong answers

Option A is wrong because a Cloud Functions + Cloud Scheduler approach requires custom code to call the Google Ads API, handle pagination, retries, and schema mapping — significantly more operational overhead than the managed connector. Option B is wrong because Cloud Composer (managed Airflow) plus Dataflow is heavyweight for a simple daily load-and-transform; it introduces DAG maintenance, cluster costs, and pipeline code that the question explicitly wants to avoid. Option C is wrong because Storage Transfer Service moves files between storage systems; it does not extract data from the Google Ads API, and CSV-based loading adds manual steps and latency.

77
MCQmedium

A financial services company receives real-time stock trade data via Pub/Sub. They need to enrich each trade with reference data from a Cloud SQL table and write the results to BigQuery for real-time analytics. The enrichment must handle late-arriving data and ensure exactly-once processing. Which Dataflow streaming pipeline configuration should be used?

A.Use a Dataflow Flex Template that reads from Pub/Sub, joins in memory, and writes to BigQuery using legacy streaming inserts
B.Use Pub/Sub to BigQuery template with streaming inserts and a side input from Cloud SQL
C.Build a custom Dataflow pipeline using the Storage Write API with exactly-once semantics and a side input from Cloud SQL
D.Deploy a Dataproc Spark Streaming job that reads from Pub/Sub, enriches via JDBC, and writes to BigQuery
AnswerC

The Storage Write API's exactly-once mode prevents duplicate writes to BigQuery, while a side input broadcasts the Cloud SQL reference table to workers for enrichment. Event-time windows with allowed lateness handle late-arriving trades, satisfying both stated requirements.

Why this answer

The BigQuery Storage Write API with exactly-once semantics is the only option that satisfies both the exactly-once requirement and the need for enrichment with Cloud SQL reference data. A custom Dataflow pipeline can use the Storage Write API's exactly-once mode while applying a side input (or CoGroupByKey) to join the streaming trades with the Cloud SQL reference table. This combination handles late-arriving data via allowed lateness and windowing while guaranteeing no duplicate writes to BigQuery.

Exam trap

PDE often tests the distinction between legacy streaming inserts (at-least-once, duplicates possible) and the Storage Write API's exactly-once mode — candidates who overlook the exactly-once requirement pick a template or legacy insert option.

How to eliminate wrong answers

Option A is wrong because legacy streaming inserts into BigQuery do not provide exactly-once semantics and can produce duplicate rows on retries. Option B is wrong because the Pub/Sub to BigQuery template uses streaming inserts (at-least-once) and does not natively support a Cloud SQL side input for enrichment. Option D is wrong because Dataproc Spark Streaming with JDBC enrichment adds operational complexity and does not inherently provide exactly-once writes to BigQuery, making it a poor fit for the stated requirements.

78
Multi-Selectmedium

A data engineer is building a Dataflow pipeline that reads newline-delimited JSON files from a Cloud Storage bucket and writes them to BigQuery. The files arrive continuously, some with malformed records, and the team wants the pipeline to keep running while routing bad records to a dead-letter location for later analysis. The team also wants the schema to be inferred from the files during development but fixed in production. Which two approaches should the data engineer use? (Choose two.)

Select 2 answers
A.Define an explicit TableSchema in the pipeline for production and use schema autodetect only in a development branch.
B.Enable the --enableStreamingEngine flag so that schema inference happens automatically at the service level.
C.Use BigQueryIO with withSchemaUpdateOptions set to ALLOW_FIELD_ADDITION and rely on runtime inference for every file.
D.Configure BigQueryIO with withCreateDisposition set to CREATE_NEVER and withWriteDisposition set to WRITE_TRUNCATE.
E.Use a DoFn with a try/catch that emits failed parses to a side output tagged as dead-letter data.
AnswersA, E

Fixing the schema explicitly in production makes writes deterministic and prevents an unexpected field in one file from altering the table layout, while schema autodetect remains useful in development for discovering the shape of new data sources. This directly satisfies the requirement that inference be used only during development while production uses a fixed schema.

Why this answer

Dead-letter routing is implemented with a side output from a DoFn that catches parse failures, keeping the main pipeline alive. Schema stability in production is achieved by declaring an explicit TableSchema rather than letting autodetect run continuously, with autodetect reserved for development. Together these two choices satisfy both the resilience and schema-governance requirements without changing write dispositions or relying on service-level flags.

Exam trap

The trap here is conflating schema inference with pipeline execution settings, when inference is a development-time client concern and dead-lettering is a transform-level pattern.

79
Multi-Selecthard

A data platform team uses Cloud Data Fusion to move data from an on-premises relational database into BigQuery. They need the pipeline to run on a fixed schedule, capture only rows changed since the last successful run, and avoid re-reading the entire source table each night. The source table has an updated_at column that is reliably populated. Which two approaches should they use? (Choose two.)

Select 2 answers
A.Use a Replicator pipeline with the default full-table snapshot mode for each run
B.Enable the Cloud Data Fusion lineage feature to track which rows were previously loaded
C.Configure the pipeline to run on a schedule using the Data Fusion scheduler or a Cloud Composer trigger
D.Use the BigQuery sink plugin with write disposition set to truncate before each nightly load
E.Use the Database source plugin with a query that filters on updated_at greater than the last recorded watermark
AnswersC, E

Cloud Data Fusion pipelines can be scheduled directly through the built-in scheduler, or triggered externally by Cloud Composer, which provides the fixed nightly cadence the team requires. Without a schedule the pipeline would have to be started manually, so this is necessary to meet the recurring-run requirement while keeping the incremental logic in the pipeline itself.

Why this answer

Scheduling the pipeline and filtering the source query on a persisted updated_at watermark together deliver a recurring job that reads only changed rows. The watermark must be written after a successful load so failures do not advance it prematurely, and the query-based Database source plugin is what makes the filter possible without custom code.

Exam trap

The trap here is treating lineage or truncate-and-reload as change-capture mechanisms, when only a stored watermark plus a filtered source query reads just the changed rows.

80
MCQmedium

A media company streams real-time viewer data from Pub/Sub to BigQuery using a Dataflow pipeline. They need to handle occasional malformed messages without losing valid data. Which pattern should they implement?

A.Raise an exception in the pipeline and stop processing
B.Use retry logic in the pipeline to reprocess malformed messages indefinitely
C.Implement a dead letter sink to store malformed messages for later analysis
D.Discard malformed messages and log an error
AnswerC

A dead letter sink diverts messages that fail parsing or validation into a separate store, so the pipeline continues processing valid records without data loss. This satisfies the requirement to handle occasional malformed messages while preserving valid viewer data for later analysis.

Why this answer

A dead letter sink (e.g., a separate Pub/Sub topic or a BigQuery error table) allows the Dataflow pipeline to route malformed messages out of the main processing path while continuing to process valid data. This pattern ensures no valid data is lost and provides a durable location for later analysis or reprocessing of the malformed records, which is essential for streaming pipelines where data quality issues are intermittent.

Exam trap

Google Cloud often tests the dead letter pattern to see if candidates understand that streaming pipelines must handle bad data gracefully without stopping or losing valid records, and the trap is that many candidates choose retry logic (Option B) because they confuse transient errors with permanent data quality issues.

How to eliminate wrong answers

Option A is wrong because raising an exception and stopping the pipeline would cause all processing to halt, leading to data loss for valid messages and violating the requirement to handle malformed messages without losing valid data. Option B is wrong because retrying malformed messages indefinitely would cause the pipeline to stall on bad records, potentially blocking the processing of subsequent valid messages and increasing latency; Dataflow's retry mechanisms are intended for transient errors, not for permanently malformed data. Option D is wrong because discarding malformed messages and logging an error results in permanent data loss, which contradicts the requirement to preserve data for later analysis and violates best practices for data integrity in streaming pipelines.

81
MCQmedium

You are designing a Dataflow pipeline to process streaming data. The pipeline may encounter malformed records. You need to handle these errors without failing the entire pipeline and store the bad records for later analysis. What is the best practice?

A.Use a dead letter sink to write malformed records to a separate Pub/Sub topic or GCS location.
B.Catch the exception and log it, then continue processing.
C.Write all records to BigQuery using the Storage Write API and handle errors in the write operation.
D.Raise an exception in the DoFn to stop the pipeline for manual intervention.
AnswerA

A dead letter sink isolates malformed records by routing them to a separate Pub/Sub topic or GCS location, so a single bad record cannot fail the whole streaming pipeline. This satisfies the requirement to keep processing valid data while retaining failures for later analysis.

Why this answer

A dead letter sink writes malformed records to a separate Pub/Sub topic or GCS location so they can be analyzed later without failing the pipeline, which is the recommended Apache Beam pattern for handling bad records. This preserves the main pipeline's throughput and keeps the bad data available for debugging or reprocessing.

Exam trap

The trap is choosing 'catch and log' as a simpler alternative, but logging does not provide a durable, queryable store for bad records, which the question explicitly requires.

How to eliminate wrong answers

Option B is wrong because catching and logging the exception loses the record content in a structured way and does not provide a durable store for later analysis; logs are not a substitute for a dead letter sink. Option C is wrong because writing all records to BigQuery and handling errors in the write operation does not isolate malformed records and may still fail the pipeline or lose data. Option D is wrong because raising an exception stops the pipeline, which directly violates the requirement to not fail the entire pipeline.

82
MCQhard

A media analytics team runs a Dataflow streaming pipeline that reads click events from Pub/Sub and writes aggregates to BigQuery. During peak hours, the pipeline's BigQuery write throughput plateaus and Dataflow logs show repeated quota-related retries on the streaming insert API. The team wants to keep exactly-once semantics and increase sustained write throughput. What should they change?

A.Enable autoscaling on the Dataflow job and raise the maximum worker count to the regional limit.
B.Write the aggregates to Cloud Storage as sharded files and schedule a load job every five minutes.
C.Increase the number of Dataflow workers to spread the streaming insert load across more threads.
D.Switch the BigQuery sink to the Storage Write API with exactly-once semantics enabled.
AnswerD

The Storage Write API is designed for high-throughput streaming ingestion and, when the exactly-once stream mode is used, provides exactly-once delivery into BigQuery. It avoids the per-row streaming insert quota that causes the plateau while preserving the semantics the team requires. This is the recommended replacement for the legacy streaming insert path in Dataflow.

Why this answer

The Storage Write API with exactly-once mode is the intended high-throughput sink for Dataflow to BigQuery and removes the legacy streaming insert quota bottleneck while keeping exactly-once delivery. Increasing workers or autoscaling addresses compute, not the API quota, and batching to Cloud Storage trades away latency and complicates semantics. Replacing the sink is the targeted fix.

Exam trap

The trap here is treating a BigQuery API quota plateau as a Dataflow scaling problem and adding workers instead of changing the write path.

83
MCQeasy

A company wants to migrate 500 TB of on-premises archival data to Cloud Storage. The data is stored on a SAN and the network link is limited to 1 Gbps. The migration must complete within 10 days. What is the MOST cost-effective approach?

A.Set up a Cloud VPN and use rsync over the encrypted connection.
B.Use BigQuery Data Transfer Service to load the data directly into BigQuery.
C.Order a Transfer Appliance, copy data locally, and ship it to Google for ingestion.
D.Use Storage Transfer Service to copy data from on-premises to GCS over the existing network.
AnswerC

A 1 Gbps link transfers roughly 10 TB per day, so 500 TB needs about 50 days — far beyond the 10-day deadline. Transfer Appliance ships data physically, bypassing the bandwidth constraint, and is cheaper than upgrading the network link.

Why this answer

The Transfer Appliance is designed for large-scale data migrations where network bandwidth is insufficient. With 500 TB at 1 Gbps, the theoretical transfer time is over 46 days, far exceeding the 10-day window. The appliance allows you to physically ship the data, bypassing network constraints entirely, making it the most cost-effective and timely solution.

Exam trap

The trap here is that candidates underestimate the time required for network transfer at 1 Gbps and overestimate the practicality of compression or incremental sync, failing to recognize that physical shipping is the only viable option for multi-petabyte data within a tight deadline.

How to eliminate wrong answers

Option A is wrong because rsync over a 1 Gbps Cloud VPN would take approximately 46 days for 500 TB (assuming full utilization, which is unrealistic due to overhead and encryption), far exceeding the 10-day deadline. Option B is wrong because BigQuery Data Transfer Service is for loading data from SaaS applications (e.g., Google Ads, Amazon S3) or other cloud sources into BigQuery, not for ingesting on-premises archival data into Cloud Storage. Option D is wrong because Storage Transfer Service relies on the existing 1 Gbps network link, which would require over 46 days for 500 TB, violating the 10-day requirement.

84
MCQhard

A data engineer is using Spark on Dataproc to process a large dataset. They notice the job is slow due to excessive shuffling. They want to optimize the job by using a more efficient data structure that reduces serialization overhead and provides better memory management. Which Spark API should they use?

A.Spark SQL
B.Spark Streaming
C.RDDs
D.DataFrames or Datasets
AnswerD

DataFrames and Datasets use Catalyst's optimised binary representation, bypassing Java serialisation and enabling off-heap Tungsten memory management. This directly reduces the serialisation overhead and excessive shuffling described in the stem, unlike RDDs, which serialise objects individually.

Why this answer

DataFrames and Datasets are built on Spark SQL's Catalyst optimizer and Tungsten execution engine, which provide schema-aware encoding and off-heap memory management. This dramatically reduces serialization overhead compared to RDDs and enables optimizations like predicate pushdown and shuffle partitioning improvements, directly addressing excessive shuffling.

Exam trap

PDE often tests the misconception that RDDs are always faster because they are 'lower level' — in fact, DataFrames/Datasets win on shuffle-heavy workloads due to Catalyst and Tungsten.

How to eliminate wrong answers

Option A is wrong because Spark SQL is the query interface/language layer, not the data structure — while DataFrames are accessed via Spark SQL, the question asks for the API/data structure that reduces serialization and improves memory management. Option B is wrong because Spark Streaming is for continuous data processing, not for optimizing shuffle-heavy batch jobs. Option C is wrong because RDDs are the low-level, untyped API that rely on Java/Kryo serialization for every shuffle and lack Catalyst/Tungsten optimizations — they are the source of the inefficiency, not the fix.

85
MCQhard

A company is using Pub/Sub to ingest clickstream events and Dataflow to write to BigQuery. They observe that some events are malformed and cause the pipeline to fail. They need a solution that captures malformed events without blocking the pipeline and allows reprocessing later. Which Dataflow pattern should they implement?

A.Use a side input to filter malformed events before the main pipeline
B.Use the Reshuffle transform to reattempt failures
C.Write malformed events to a dead letter sink (e.g., another Pub/Sub topic or GCS bucket) and continue processing healthy events
D.Use logging alerts to notify the team and stop the pipeline on error
AnswerC

Routing malformed records to a dead letter sink isolates them from the main pipeline, so healthy events continue flowing to BigQuery. The sink preserves the bad records for later inspection and reprocessing, satisfying the non-blocking capture requirement.

Why this answer

The dead letter sink pattern is the canonical Dataflow approach for handling malformed or unparseable records: instead of throwing an exception that fails the bundle (and potentially the whole pipeline), the transform routes the bad record to a secondary sink such as a Pub/Sub topic or GCS bucket while healthy events continue downstream. This preserves pipeline throughput and keeps the malformed payloads available for later inspection, correction, and reprocessing. It is the standard implementation of the dead-letter queue pattern in Apache Beam.

Exam trap

PDE often tests the misconception that side inputs or Reshuffle can handle bad records — candidates confuse data distribution/auxiliary data mechanisms with error-handling patterns, when only a dead-letter sink actually isolates failures.

How to eliminate wrong answers

Option A is wrong because a side input is a read-only auxiliary dataset broadcast to a transform (e.g., a lookup table), not a mechanism for isolating and persisting failed records. Option B is wrong because Reshuffle only redistributes elements across workers to break fusion or improve parallelism; it does not catch exceptions or retry failed elements. Option D is wrong because stopping the pipeline on error is the exact opposite of the requirement — it blocks processing of healthy events and provides no reprocessing path.

86
MCQhard

A streaming pipeline ingests events from Pub/Sub, enriches them via a slow REST API call, and writes the result to BigQuery. The API has a limit of 10 requests per second per client. The pipeline processes 1000 messages per second. Which approach minimizes latency while respecting API limits?

A.Use a global window with a trigger that fires every second, and inside the DoFn limit concurrent API calls to 10.
B.Fan out the stream to multiple REST API instances using Pub/Sub topic splitting.
C.Use a Dataflow Flex Template to run multiple pipelines, each processing a subset of messages.
D.Assign each message a random key and use a sliding window of 10 seconds; the API call will be distributed across workers.
AnswerA

Groups messages into batches per second, then controls concurrency to stay within the 10 req/s limit.

Why this answer

Using a global window with a trigger every second groups 1000 messages into a batch, and then throttling concurrent API calls to 10 within the DoFn (e.g., using a fixed-size thread pool) respects the API limit while minimizing latency by processing messages in parallel up to the limit. Option B is wrong because fanning out to multiple API instances doesn't help if the limit is per client; the total requests per second across all instances would still exceed the client limit. Option C is wrong because Dataflow Flex Templates are used to run parameterized pipelines, not to solve throttling issues.

Option D is wrong because assigning a random key and using a sliding window distributes messages across workers, but without explicit throttling, the API limit could still be exceeded.

87
MCQhard

A logistics company ingests GPS pings from delivery vans into Pub/Sub, and a Dataflow streaming pipeline writes them to BigQuery. Latency requirements are lenient (about 5 minutes), but the finance team needs the pipeline's cost to be predictable and low, and the data volume fluctuates by a factor of ten between day and night. The team wants to minimize per-element cost without losing data. Which configuration should the data engineer choose?

A.Streaming Engine with autoscaling and a bounded maximum worker count, plus windowing with a 5-minute trigger so BigQuery writes are batched.
B.Batch mode with a 5-minute micro-batch schedule that reads Pub/Sub snapshots and writes to BigQuery.
C.Streaming Engine with autoscaling, a low --maxNumWorkers, and a streaming trigger that emits results at least every 5 minutes.
D.Streaming Engine with a fixed worker count sized for peak daytime load and --maxNumWorkers capped at that value.
AnswerA

Autoscaling lets Dataflow add workers during the daytime surge and shrink at night, matching cost to load while a bounded maximum keeps spend predictable. Windowing with a five-minute trigger batches BigQuery writes, which reduces per-row streaming insert overhead and cost. This satisfies the lenient latency target while keeping finance's cost model stable and avoiding data loss.

Why this answer

The scenario combines wildly variable throughput with a lenient latency target and a hard cost ceiling. Autoscaling with a bounded maximum worker count is the standard way to match capacity to demand while capping worst-case spend, and a five-minute windowing trigger amortizes BigQuery write overhead across many elements. Together these choices honor the latency allowance, keep costs predictable, and avoid the data-loss risk of an undersized fixed pool.

Exam trap

The trap here is treating low cost as always meaning fewer workers, when the real requirement is elasticity in both directions with a bounded maximum.

88
Multi-Selecthard

Your company has a Dataproc cluster that runs Spark jobs. You need to choose between RDDs, DataFrames, and Datasets for a new job that performs complex aggregations on structured data. Which TWO statements are correct regarding performance and ease of use?

Select 2 answers
A.DataFrames and Datasets are both available in PySpark.
B.DataFrames store data in a columnar format, allowing better compression.
C.RDDs are easier to use than DataFrames for complex aggregations.
D.DataFrames are optimized by Spark's Catalyst optimizer, leading to faster execution.
E.Datasets provide compile-time type safety and are always faster than DataFrames.
AnswersB, D

DataFrames use a columnar in-memory representation with Catalyst optimisation and Tungsten encoding, yielding better compression and scan efficiency than row-based RDDs. For complex aggregations on structured data, this columnar layout reduces I/O and memory footprint, improving performance.

Why this answer

Option B is correct because Spark DataFrames are built on Tungsten's columnar in-memory representation, which stores data by column and enables more efficient encoding, compression, and cache-friendly scans than RDDs' row-based Java/Python objects. Option D is correct because DataFrame operations are analyzed and rewritten by the Catalyst optimizer, which applies rule-based and cost-based optimizations such as predicate pushdown, column pruning, and join reordering, producing faster execution plans than hand-written RDD transformations. Option A is incorrect because Datasets are not available in PySpark; they exist only in the Scala and Java APIs, while PySpark offers DataFrames (and RDDs).

Option C is incorrect because RDDs are lower-level and require manual implementation of aggregation logic, making them harder, not easier, than DataFrames for complex aggregations. Option E is incorrect because, although Datasets add compile-time type safety in Scala/Java, they are not always faster than DataFrames and often incur serialization overhead; the claim of being 'always faster' is false.

Exam trap

PDE often tests the misconception that Datasets are always faster or that they exist in PySpark — candidates who assume 'typed = better performance' or 'Python has Datasets' pick the wrong answers.

89
MCQhard

A healthcare company needs to process HL7 messages containing sensitive patient data. The messages arrive in Cloud Storage as JSON files. The pipeline must de-identify the data using the Cloud Healthcare API DLP de-identification, then load the results into BigQuery. The pipeline must ensure that no unredacted data is ever written to BigQuery, and that processing is fault-tolerant. The Dataflow pipeline reads from Cloud Storage, calls the DLP API for de-identification, and writes to BigQuery. Which additional configuration ensures that only de-identified data reaches BigQuery?

A.Use a ParDo that calls the DLP API and writes the de-identified data to BigQuery; if the DLP call fails, retry indefinitely until success.
B.Use a DoFn that calls the DLP API and only emits the de-identified record; then use a BigQueryIO write with a dead-letter queue for failed DLP calls.
C.Use a ParDo transform that calls the DLP API and then writes the original and de-identified data to separate BigQuery tables, with access controls on the original table.
D.Use a GroupByKey to batch records, then call the DLP API in a batch request; write both original and de-identified data to BigQuery with column-level encryption.
AnswerB

By only emitting de-identified records, the pipeline ensures that unredacted data never reaches BigQuery. Using a dead-letter queue for failed DLP calls prevents data loss and allows reprocessing. This design is fault-tolerant and meets the strict requirement that no original data is written to BigQuery. It also handles API errors gracefully without compromising data privacy.

Why this answer

The pipeline must ensure that only de-identified data is written to BigQuery. Emitting only de-identified records from the DoFn guarantees that unredacted data never reaches the sink. A dead-letter queue for failed DLP calls maintains fault tolerance by isolating problematic records for later analysis or reprocessing, rather than blocking the pipeline or writing original data.

This design meets both privacy and reliability requirements.

Exam trap

The trap here is assuming that encryption or access controls on original data in BigQuery are sufficient, but the requirement explicitly forbids any unredacted data in BigQuery.

90
MCQhard

A media company ingests clickstream events into Pub/Sub and processes them with a Dataflow streaming pipeline that writes to BigQuery. The pipeline uses a fixed window of five minutes and discards late data. A product manager reports that events arriving more than five minutes after their event timestamp never appear in reports. Which change should you make to capture those events?

A.Switch the pipeline to use the Pub/Sub message publish time as the element timestamp.
B.Add a GroupByKey transform before the window to buffer all events.
C.Increase the fixed window size to thirty minutes.
D.Configure an allowed lateness on the window and adjust the trigger to emit updated results.
AnswerD

Allowed lateness extends how long a window accepts elements after its end, and a trigger that fires on late data emits revised results for those windows. Together they let events arriving beyond five minutes be incorporated into the aggregation instead of being discarded, which directly addresses the missing late arrivals.

Why this answer

Discarding late data is a consequence of the window's allowed lateness and trigger configuration, not of window duration. Setting an allowed lateness period and a trigger that emits on late arrivals lets the pipeline accept events beyond the five-minute boundary and update the affected windows, so the stragglers appear in reports.

Exam trap

The trap here is assuming that a larger window automatically captures late events, when lateness handling is governed by allowed lateness and triggers, not by window length.

91
MCQeasy

You need to stream real-time user click events from your application into BigQuery for immediate analysis. The events must be available for query within seconds. Which approach is recommended?

A.Use Pub/Sub to Dataflow to BigQuery with the Storage Write API for high-throughput streaming.
B.Use Cloud Data Fusion to ingest streaming data from Pub/Sub into BigQuery.
C.Use Cloud Functions to receive events from Pub/Sub and insert them into BigQuery using the legacy streaming API.
D.Use Pub/Sub with a BigQuery subscription to directly write events into BigQuery.
AnswerA

Pub/Sub with Dataflow and the Storage Write API delivers true streaming ingestion, satisfying the seconds-level latency constraint. Dataflow handles windowing and exactly-once processing, while the Storage Write API commits rows directly into BigQuery without the batching delays of load jobs or legacy streaming inserts.

Why this answer

The recommended approach is to use Pub/Sub to Dataflow to BigQuery with the Storage Write API. Dataflow provides a managed stream processing service that can handle high-throughput, low-latency ingestion, and the Storage Write API offers exactly-once semantics and is optimized for streaming inserts into BigQuery. This combination ensures events are available for query within seconds and scales well.

Exam trap

PDE often tests the difference between the legacy streaming API and the Storage Write API, and candidates may incorrectly choose the simpler Pub/Sub to BigQuery subscription without considering throughput and exactly-once requirements.

How to eliminate wrong answers

Option B is wrong because Cloud Data Fusion is a GUI-based data integration tool that is more suited for batch and ETL workloads, not for low-latency real-time streaming; it would introduce higher latency. Option C is wrong because using Cloud Functions to insert into BigQuery with the legacy streaming API is not recommended for high-throughput scenarios due to potential bottlenecks, lack of exactly-once semantics, and higher cost; the legacy API also has limitations. Option D is wrong because a BigQuery subscription to Pub/Sub directly writes to BigQuery but does not provide the same level of processing, transformation, and exactly-once guarantees as Dataflow with the Storage Write API; it is also limited in throughput and may not meet the 'within seconds' requirement for high volumes.

92
MCQeasy

A data engineer needs to query a BigQuery table that contains an array of structs. They want to expand the array into separate rows for each element. Which SQL function should they use?

A.STRUCT
B.UNNEST
C.ARRAY_AGG
D.SPLIT
AnswerB

UNNEST expands an array into a set of rows, one per element, which is exactly the flattening required. Placed in the FROM clause alongside the table, it correlates each struct element into its own row for querying.

Why this answer

UNNEST is the BigQuery SQL function that flattens an array into a set of rows, one per array element. When applied to an array of structs, it expands each struct into its own row, allowing the struct fields to be selected as columns.

Exam trap

The trap is confusing array construction (STRUCT, ARRAY_AGG) with array expansion (UNNEST) — candidates who see 'array of structs' sometimes reach for STRUCT or ARRAY_AGG instead of the flattening function.

How to eliminate wrong answers

Option A is wrong because STRUCT constructs a struct value; it does not expand arrays into rows. Option C is wrong because ARRAY_AGG aggregates values into an array, the opposite of what is needed. Option D is wrong because SPLIT divides a string into an array of substrings, which is unrelated to flattening arrays of structs.

93
Multi-Selectmedium

A data engineer is designing a batch processing pipeline that runs daily. The pipeline reads CSV files from GCS, transforms them using Python, and writes the results to BigQuery. They need to parameterize the pipeline for different environments and run it on a schedule. Which THREE components should they use? (Choose 3)

Select 3 answers
A.Cloud Composer
B.Dataproc
C.Dataflow Flex Template
D.Storage Transfer Service
E.Cloud Functions
AnswersA, B, C

Cloud Composer orchestrates and schedules the pipeline, and allows parameterization via Airflow variables.

Why this answer

Cloud Composer (A) is correct because it is a managed Apache Airflow service that provides DAG-based scheduling and parameterization, making it ideal for orchestrating a daily batch pipeline with environment-specific variables. Dataproc (B) is correct because it offers managed Spark/Hadoop clusters that can run Python-based transformations over the CSV data read from GCS at batch scale. Dataflow Flex Template (C) is correct because it packages a Dataflow pipeline (including Python transforms) into a reusable, parameterized template that can be invoked with different runtime parameters per environment.

Storage Transfer Service (D) is not appropriate here because it only moves data between storage systems and performs no transformation or scheduling logic. Cloud Functions (E) is unsuitable because it is an event-driven, short-lived serverless function, not designed for orchestrating daily batch pipelines or large-scale data transformation.

Exam trap

Google often tests the distinction between orchestration/scheduling services (Cloud Composer) and compute/processing services (Dataproc, Cloud Functions), leading candidates to mistakenly choose Dataproc for scheduling or Cloud Functions for batch processing.

94
Multi-Selecthard

You are building a Dataflow pipeline that reads from Pub/Sub, applies transformations, and writes to BigQuery. The pipeline must handle late-arriving data and ensure that the windowing and triggering are correct. Which THREE configurations should you consider? (Choose 3)

Select 3 answers
A.Enable Dataflow Streaming Engine for exactly-once processing.
B.Use side inputs to enrich streaming data with static data.
C.Use the BigQuery Storage Write API with committed mode to ensure exactly-once writes.
D.Set an allowed lateness duration to handle late-arriving data.
E.Configure a triggering frequency to control how often results are emitted.
AnswersC, D, E

The BigQuery Storage Write API's committed mode provides exactly-once semantics through write-stream offsets, preventing duplicate rows when retries occur after late data triggers additional pane firings. This satisfies the pipeline's requirement for correct windowing and triggering, since speculative or repeated emissions from allowed lateness would otherwise insert duplicates into BigQuery.

Why this answer

Option C is correct because the BigQuery Storage Write API in committed mode (exactly-once semantics) uses write streams with offsets and client-side deduplication, which is the recommended way to guarantee exactly-once writes from a Dataflow streaming pipeline into BigQuery. Option D is correct because setting an allowed lateness duration (via withAllowedLateness) on the windowing strategy tells Dataflow how long to retain window state and keep accepting late-arriving elements after the watermark passes the window end, which is exactly the requirement for handling late data. Option E is correct because configuring a triggering frequency (for example, AfterProcessingTime.pastFirstElementInPane().plusDelayOf(...) or a repeated trigger) controls how often intermediate or final results are emitted from each window, which directly governs the windowing/triggering behavior described in the scenario.

Option A does not belong because Dataflow Streaming Engine is a runner feature that offloads state and windowing execution to the service for scalability and cost benefits; it does not itself provide exactly-once processing, which comes from the pipeline's sources, sinks, and deduplication logic. Option B does not belong because side inputs are used to enrich streaming records with additional (often static or slowly changing) data, which is unrelated to handling late-arriving data or to windowing and triggering correctness.

Exam trap

A common misconception is that Dataflow Streaming Engine alone provides exactly-once processing. In reality, it is the combination of source/sink semantics (like the BigQuery Storage Write API) that ensures exactly-once, not the engine itself.

95
Multi-Selectmedium

A company needs to stream real-time user activity data from their application into BigQuery for immediate dashboarding. They want to minimize latency (under 5 seconds) and ensure exactly-once delivery. Which TWO options should they consider? (Choose 2)

Select 2 answers
A.Use Cloud Functions to receive events and call the BigQuery REST API
B.Use BigQuery Storage Write API in committed mode
C.Use BigQuery legacy streaming inserts directly from the application
D.Use Apache Kafka on Dataproc and write to BigQuery via the BigQuery Kafka connector
E.Stream data to Pub/Sub, then use Dataflow to write to BigQuery with exactly-once guarantees
AnswersB, E

BigQuery Storage Write API in committed mode provides exactly-once semantics through stream-level offsets, so duplicate records are discarded on retry. It supports sub-second, real-time ingestion directly into BigQuery, satisfying the under-five-second latency requirement without intermediate staging. This makes it ideal for streaming live user activity into dashboards.

Why this answer

Option B is correct because the BigQuery Storage Write API in committed mode provides exactly-once semantics via write streams and offsets, and it supports low-latency streaming ingestion suitable for sub-5-second dashboarding. Option E is correct because Pub/Sub plus Dataflow is the canonical Google Cloud streaming pipeline: Pub/Sub ingests events durably, and Dataflow's BigQueryIO in STREAMING mode with exactly-once processing guarantees deduplication and exactly-once writes to BigQuery. Option A is not ideal because Cloud Functions invoking the BigQuery REST API (tabledata.insertAll) does not provide exactly-once delivery and adds per-event overhead and cold-start latency.

Option C is wrong because legacy streaming inserts offer at-least-once semantics and can produce duplicate rows, violating the exactly-once requirement. Option D is not the best fit because running Kafka on Dataproc adds operational complexity, and the Kafka connector typically relies on the Storage Write API or streaming inserts without inherently guaranteeing exactly-once end-to-end delivery in this scenario.

Exam trap

Google Cloud often tests the distinction between legacy streaming inserts (at-least-once) and the Storage Write API in committed mode (exactly-once). Candidates may mistakenly choose legacy inserts because they are simpler to implement, ignoring the exactly-once requirement.

96
MCQeasy

A retail company wants to analyze point-of-sale transaction data stored in Cloud SQL for PostgreSQL. They need to run complex analytical queries joining this data with data in BigQuery. The data changes frequently, and they want near-real-time access without impacting the production Cloud SQL instance. Which approach should they use?

A.Export Cloud SQL data to Cloud Storage daily and load into BigQuery
B.Use BigQuery federated queries to query Cloud SQL directly
C.Use Cloud Data Fusion to replicate Cloud SQL to BigQuery with a batch pipeline
D.Set up a Datastream stream from Cloud SQL to BigQuery
AnswerD

Datastream provides change data capture (CDC) from Cloud SQL for PostgreSQL to BigQuery, replicating changes in near-real-time. It reads from the source's replication log without impacting production query performance and writes to BigQuery with low latency. This enables complex analytical joins in BigQuery while keeping the production database isolated.

Why this answer

Datastream is a serverless change data capture service that replicates data from Cloud SQL for PostgreSQL to BigQuery in near-real-time without impacting the source database's performance. It reads the write-ahead log and streams changes, enabling up-to-date analytics. Other options either introduce latency, impact production, or require manual intervention.

Exam trap

The trap here is choosing federated queries because they seem to provide direct access, but they can impact production and do not synchronize data continuously.

97
MCQmedium

A logistics company uploads millions of small JSON files per day to a Cloud Storage bucket and needs to query them with standard SQL immediately, but the analytics team does not want to create BigQuery tables or manage schema changes. The files follow a consistent structure but new fields are added frequently. Which BigQuery capability should they use to query these objects directly from Cloud Storage with the least operational overhead?

A.Use the BigQuery Storage Write API to stream the JSON files into a table.
B.Create an external table over the Cloud Storage bucket and query it with standard SQL.
C.Mount the Cloud Storage bucket as a BigQuery dataset and query the files as tables.
D.Load the JSON files into a BigQuery table using a recurring Dataflow batch job.
AnswerB

BigQuery external tables let you query data stored in Cloud Storage with standard SQL without loading it, and you can define the schema or let BigQuery infer it. This matches the requirement to avoid managing loads and table schema changes, though query performance is lower than native tables and you are billed for bytes scanned in Cloud Storage.

Why this answer

An external table defined over Cloud Storage allows standard SQL queries against the JSON objects without loading them, and schema autodetection handles evolving fields. This minimizes operational overhead because there is no pipeline or table to maintain, at the cost of query performance and per-byte scanning charges on the external data.

Exam trap

The trap here is assuming that querying Cloud Storage data requires loading it into BigQuery first; external tables exist precisely to query objects in place.

98
MCQeasy

A data engineer needs to transfer 500 TB of on-premises data to Google Cloud Storage. The data is stored on NAS devices and the network bandwidth is limited to 100 Mbps. What is the most cost-effective and timely transfer method?

A.Use Storage Transfer Service over the internet
B.Use a VPN connection and rsync
C.Use gsutil cp in parallel
D.Use Transfer Appliance
AnswerD

Transfer Appliance bypasses the 100 Mbps network constraint entirely by shipping physical hardware, making 500 TB feasible in days rather than the months a 100 Mbps link would require. For NAS-resident data at this volume, it is also more cost-effective than sustained egress or Interconnect charges.

Why this answer

Transfer Appliance is Google's physical data transfer service that ships a rackable appliance to the customer site, where data is copied locally and the appliance is shipped back to Google for upload to Cloud Storage. For 500 TB over a 100 Mbps link, the network transfer time would be roughly 1.5 years, making online transfer impractical. Transfer Appliance is the most cost-effective and timely option for this volume and bandwidth.

Exam trap

The trap is underestimating transfer time — candidates pick Storage Transfer Service or gsutil because they are 'cloud-native,' forgetting that 500 TB over 100 Mbps is physically impossible in a reasonable timeframe.

How to eliminate wrong answers

Option A is wrong because Storage Transfer Service over the internet is still bounded by the 100 Mbps link, so 500 TB would take over a year and incur significant egress/transfer costs. Option B is wrong because a VPN plus rsync is even slower due to encryption overhead and is not designed for petabyte-scale bulk migration. Option C is wrong because gsutil cp in parallel only improves throughput marginally and cannot overcome a 100 Mbps physical link limit.

99
MCQmedium

A data engineer is using Apache Spark on Dataproc to process a large dataset. They need to perform complex aggregation and transformation with high performance. The dataset has a known schema and they want to take advantage of Catalyst optimizer. Which Spark API should they use?

A.Spark SQL only
B.DataFrames
C.Datasets
D.RDDs
AnswerB

DataFrames expose a structured schema, letting Spark's Catalyst optimizer apply rule-based and cost-based query planning to aggregations and transformations. This delivers the high performance the scenario demands, unlike RDDs, which lack schema awareness and bypass Catalyst entirely.

Why this answer

DataFrames are the correct choice because they combine a declarative, structured API with the Catalyst optimizer, which builds and optimizes a logical plan before execution. With a known schema, DataFrames let Spark apply rule-based and cost-based optimizations such as predicate pushdown, column pruning, and join reordering. This delivers high performance for complex aggregations and transformations without sacrificing the structured API benefits.

Exam trap

The trap here is confusing Spark SQL (a query interface) with DataFrames (the structured API that actually invokes Catalyst), or assuming Datasets are always the best choice when language support and schema handling matter.

How to eliminate wrong answers

Option A is wrong because Spark SQL is a subset/interface for SQL queries and does not by itself provide the full programmatic DataFrame API needed for complex transformations; it is not the broadest answer. Option C is wrong because Datasets are a typed extension available primarily in Scala/Java and are not the general-purpose API for all languages; the question asks for the API that leverages Catalyst with a known schema, which DataFrames do. Option D is wrong because RDDs are the low-level, unstructured API that bypasses Catalyst and Tungsten optimizations, so they cannot take advantage of the optimizer.

100
MCQmedium

A data engineer needs to create a Dataflow pipeline template that can be reused across multiple environments (dev, staging, prod) with different parameters (e.g., input Pub/Sub topic, output BigQuery table). Which template type should they use?

A.Dataflow Prime
B.Flex Template
C.Classic Template
D.Cloud Composer workflow template
AnswerB

Flex Templates package the pipeline as a Docker image with runtime parameters supplied via metadata, so the same template runs across dev, staging and prod by passing different Pub/Sub topics and BigQuery tables. Classic templates require recompilation for such changes.

Why this answer

Flex Templates (B) are the correct choice because they package a Dataflow pipeline as a Docker image, allowing environment-specific parameters (e.g., Pub/Sub topic, BigQuery table) to be passed at runtime via the --parameters flag. This enables true reusability across dev, staging, and prod without modifying the template code, unlike Classic Templates which require compile-time parameterization.

Exam trap

The Google Cloud exam often tests the distinction between Classic Templates (compile-time parameterization) and Flex Templates (runtime parameterization), trapping candidates who assume all templates support the same level of parameter flexibility.

How to eliminate wrong answers

Option A is wrong because Dataflow Prime is a managed service for optimizing resource utilization and autoscaling, not a template type for parameterized reuse. Option C is wrong because Classic Templates require parameters to be baked in at staging time, making them less flexible for multi-environment reuse without rebuilding the template. Option D is wrong because Cloud Composer is an Apache Airflow orchestration service used to schedule and monitor workflows, not a Dataflow template type for parameterized pipeline reuse.

101
MCQhard

You are designing a streaming pipeline that must handle late-arriving data with a maximum lateness of 10 minutes. You need to ensure that all data is processed exactly once and that results are emitted after the watermark passes the window. Which Apache Beam concept should you use to achieve this?

A.Use sliding windows with a period of 10 minutes and a trigger that fires when the watermark passes the end of the window.
B.Use a global window with a trigger that fires every 10 minutes.
C.Use session windows with a gap duration of 10 minutes and a trigger that fires on every element.
D.Use fixed windows with allowed lateness set to 10 minutes and a trigger that fires when the watermark passes the end of the window.
AnswerD

Setting allowed lateness to 10 minutes on fixed windows ensures that late data up to 10 minutes is not dropped. A trigger that fires when the watermark passes the end of the window emits results only after the watermark indicates all on-time data has arrived, satisfying the requirement.

Why this answer

Fixed windows with allowed lateness of 10 minutes and a watermark-based trigger ensure that data arriving up to 10 minutes late is included, and results are emitted only after the watermark passes the window end, meaning all on-time data has arrived. This combination provides exactly-once processing semantics when used with a runner that supports it.

Exam trap

The trap here is confusing processing-time triggers with event-time watermarks; only watermark-based triggers respect event-time completeness and allowed lateness.

102
MCQhard

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. Some incoming messages are malformed and fail to parse. How should you handle these messages to ensure the pipeline continues processing without data loss?

A.Configure Pub/Sub to retry indefinitely until the message is processed
B.Use a try-catch block in the pipeline and ignore malformed messages
C.Write malformed messages to a dead-letter sink (e.g., Pub/Sub topic or GCS) and continue processing
D.Set the pipeline to fail and alert the team via Cloud Monitoring
AnswerC

Routing unparseable records to a dead-letter sink isolates them from the main pipeline, so a single malformed message cannot halt processing or be silently dropped. This satisfies the no-data-loss constraint by preserving failures for later inspection while valid messages continue to BigQuery.

Why this answer

Writing malformed messages to a dead-letter sink (Pub/Sub topic or GCS) preserves the data for later inspection while allowing the pipeline to continue processing valid messages. This is the standard Dataflow/Apache Beam pattern for handling unparseable records without halting the stream or silently discarding data. It satisfies both the 'no data loss' and 'continue processing' requirements.

Exam trap

PDE often tests the misconception that Pub/Sub retry or pipeline failure is an acceptable error-handling strategy — candidates forget that retries cannot fix deterministic parse failures and that failing the pipeline violates availability requirements.

How to eliminate wrong answers

Option A is wrong because Pub/Sub itself cannot parse or validate message content — retrying indefinitely just redelivers the same malformed payload forever, causing a poison-pill loop and blocking the subscription. Option B is wrong because silently ignoring malformed messages causes data loss and removes any audit trail for troubleshooting. Option D is wrong because failing the entire pipeline halts processing of all valid messages, violating the requirement that the pipeline continue processing.

103
MCQeasy

A data engineer needs to schedule recurring nightly loads from Amazon S3 to Google Cloud Storage. The data is in CSV format and the volume is approximately 500 GB per night. Which Google Cloud service should they use?

A.Transfer Appliance
B.Storage Transfer Service
C.BigQuery Data Transfer Service
D.Datastream
AnswerB

Storage Transfer Service performs server-side, managed transfers between object stores, including S3 to Cloud Storage, and handles scheduling, retries and 500 GB nightly volumes without compute provisioning. CSV format is irrelevant since it copies objects verbatim.

Why this answer

The Storage Transfer Service is designed for online data transfers from external cloud providers like Amazon S3 to Google Cloud Storage. It supports scheduling recurring nightly transfers, handles large volumes (500 GB/night), and automatically retries failed transfers, making it the correct choice for this use case.

Exam trap

Google often tests the distinction between services that transfer data to Cloud Storage (Storage Transfer Service) versus services that load data directly into BigQuery (BigQuery Data Transfer Service), causing candidates to confuse the destination.

How to eliminate wrong answers

Option A is wrong because Transfer Appliance is a physical device for offline data transfer, used for very large datasets (hundreds of TB to PB) where network transfer is impractical, not for recurring nightly loads. Option C is wrong because BigQuery Data Transfer Service is for loading data into BigQuery tables from sources like Google Ads or Amazon S3, but it does not directly transfer files to Cloud Storage; it loads data into BigQuery, not Cloud Storage. Option D is wrong because Datastream is for real-time change data capture (CDC) from databases like MySQL or PostgreSQL to BigQuery or Cloud Storage, not for batch CSV file transfers from S3.

104
MCQhard

A media company streams user interaction events into Pub/Sub and processes them with a Dataflow streaming pipeline that writes to BigQuery. During peak hours, the pipeline's watermark lags significantly behind real time, and late-arriving events are being dropped. The team wants late events to be included in windowed aggregations for up to 30 minutes after the window closes. Which Dataflow configuration should they apply?

A.Switch the pipeline to use processing-time windows instead of event-time windows.
B.Increase the number of worker threads and enable autoscaling to reduce the watermark lag.
C.Set the allowed lateness on the windowing transform to 30 minutes and update the aggregation to emit late panes.
D.Configure the Pub/Sub subscription to retain messages for 30 minutes and replay them.
AnswerC

Allowed lateness controls how long after a window closes Dataflow continues to accept and process late-arriving elements for that window. Setting it to 30 minutes and ensuring the aggregation emits late panes lets events arriving within that period update results, which directly addresses the dropped late events while keeping the watermark behavior unchanged.

Why this answer

Allowed lateness extends the period during which a window accepts late data after the watermark passes the window end. Setting it to 30 minutes and emitting late panes ensures events arriving within that window update the aggregation instead of being dropped, which is the correct way to handle late-arriving streaming data.

Exam trap

The trap here is confusing watermark lag with allowed lateness; scaling workers addresses throughput, not the windowing rule that discards late elements.

105
MCQhard

A financial company uses a Dataflow streaming pipeline to read transactions from Pub/Sub and write to BigQuery. They need exactly-once processing semantics for the BigQuery writes and want to avoid duplicates during pipeline updates. Which approach should they use?

A.Use the BigQuery Storage Write API with exactly-once semantics in the Dataflow BigQueryIO connector.
B.Use BigQueryIO with STREAMING_INSERTS and enable insertId-based deduplication.
C.Use BigQueryIO with the FILE_LOADS method and trigger frequent load jobs.
D.Write to a temporary table and run a MERGE statement periodically to deduplicate.
AnswerA

The BigQueryIO connector can use the Storage Write API with exactly-once semantics, which deduplicates writes using stream offsets and supports seamless pipeline updates without duplicates. This is the recommended approach for exactly-once streaming writes to BigQuery in Dataflow.

Why this answer

The Storage Write API with exactly-once semantics in BigQueryIO provides deduplication through stream offsets and supports pipeline updates without duplicates. It is the designed solution for exactly-once streaming writes to BigQuery in Dataflow, unlike legacy streaming inserts or file loads.

Exam trap

The trap here is assuming that insertId deduplication with streaming inserts guarantees exactly-once, when it only provides best-effort deduplication.

106
Multi-Selecthard

A company uses Workflows to orchestrate a multi-step data pipeline. One step calls an HTTP endpoint that may take up to 10 minutes, but the default Workflows timeout is too short. They also need to handle transient errors with retries. Which TWO configurations should they apply? (Choose 2)

Select 2 answers
A.Set a step timeout of 600 seconds for the HTTP call step
B.Configure a dead letter queue for failed steps
C.Use the default retry policy on the step
D.Set the workflow execution timeout to 600 seconds
E.Add a retry policy on the step with appropriate conditions for transient errors
AnswersA, E

The HTTP call can run for 10 minutes, exceeding the default step timeout. Setting a 600-second step timeout on that call allows the long-running request to finish rather than being cancelled mid-flight, directly satisfying the stem's stated duration constraint.

Why this answer

Option A is correct because Workflows steps support a per-step timeout, and setting it to 600 seconds (10 minutes) ensures the HTTP call step is allowed to run for the full duration instead of being cut off by the shorter default step timeout. Option E is correct because transient errors (such as 5xx responses or connection resets) should be handled by attaching a retry policy to the step with conditions that match those transient failures, so the step is retried automatically. Option B is not appropriate because Workflows does not use dead letter queues for failed steps; failures are surfaced through execution errors and logs.

Option C is wrong because relying on the default retry policy does not let them target transient errors with appropriate conditions, and the default may not retry at all. Option D is wrong because setting the workflow execution timeout to 600 seconds would cap the entire workflow at 10 minutes, which is not the same as giving the HTTP step enough time and could still cause premature termination.

107
MCQmedium

A media analytics company ingests clickstream events into Pub/Sub at a sustained rate of 2 GB/s. A Dataflow streaming pipeline reads these events, performs windowed aggregations, and writes results to BigQuery. The pipeline must handle occasional spikes up to 5 GB/s without data loss or excessive backlog. The operations team wants to minimize manual intervention and cost. What should you do to configure the Dataflow pipeline for dynamic scaling?

A.Set the number of workers to a fixed value of 100 to handle peak load, and enable autoscaling with a maximum of 100 workers.
B.Enable autoscaling and set the maximum number of workers to 200, and use Streaming Engine to offload windowing and state management.
C.Configure the pipeline to use a single worker with a high-memory machine type to reduce coordination overhead.
D.Use a batch pipeline with a trigger that runs every 5 minutes to process accumulated Pub/Sub messages.
AnswerB

Enabling autoscaling allows Dataflow to add workers when backlog increases and remove them when load decreases, optimizing cost. Setting a high maximum ensures capacity for spikes. Streaming Engine moves pipeline state and windowing out of worker memory, improving scalability and reducing worker resource needs. This combination handles dynamic load without manual intervention and is the recommended approach for variable streaming workloads.

Why this answer

Autoscaling with a sufficient maximum worker count allows Dataflow to dynamically adjust to load, while Streaming Engine optimizes state management and reduces worker burden. This combination handles both sustained high throughput and spikes without manual tuning, and it minimizes cost during low periods. Fixed worker counts or batch processing would either be inefficient or fail to meet latency and scalability needs.

Exam trap

The trap here is assuming that setting a high fixed number of workers is sufficient for spikes, but that ignores cost optimization and the benefits of dynamic scaling.

← PreviousPage 2 of 2 · 107 questions total

Ready to test yourself?

Try a timed practice session using only Ingesting and Processing the Data questions.