Courseiva

PDE · domain

Ingesting and Processing the Data

This domain covers moving data into and through Google Cloud: Pub/Sub ingestion, Dataflow/Apache Beam pipelines, Cloud Storage event triggers, BigQuery loading, Dataproc, and transfer options like Storage Transfer Service and Transfer Appliance. Questions are scenario-based, asking you to pick the right service, handle failures such as malformed records, and design pipelines that keep running rather than aborting.

107 questions31 easy45 medium31 hard

Focused practice

Practice Ingesting and Processing the Data questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Ingesting and Processing the Data

Be able to design ingestion and processing pipelines that keep running despite bad data, and choose the correct Google Cloud service for each source and trigger. The most important thing: route malformed records to a dead-letter or side output rather than letting them fail the pipeline.

Using Dataflow dead-letter patterns and side outputs to route malformed Pub/Sub or JSON records without failing the pipeline

Choosing Pub/Sub, Eventarc, or Cloud Storage notifications to trigger Cloud Run or Cloud Functions on object events

Selecting Storage Transfer Service versus Transfer Appliance for large on-premises Hadoop-to-Cloud Storage migrations

Building Apache Beam pipelines that read Cloud Storage, transform, and write to BigQuery with correct windowing and error handling

Watch out for

Common Ingesting and Processing the Data exam traps

  • ▸Letting a single malformed record throw an exception that kills the whole Dataflow pipeline instead of routing bad records to a dead-letter sink or side output.
  • ▸Assuming Pub/Sub or Cloud Storage triggers automatically invoke Cloud Run without configuring Eventarc or notifications and the required IAM permissions.
  • ▸Picking Transfer Appliance for a transfer that fits within network bandwidth and deadline, when Storage Transfer Service over the network would suffice.

Question index

All Ingesting and Processing the Data questions (107)

Click any question to see the full explanation, or start a practice session above.

1

A data engineer needs to transfer 5 PB of historical data from an on-premises Hadoop cluster to Cloud Storage. The network bandwidth is limited to 1 Gbps, and the transfer must complete within 30 days. Which transfer method should they use?

Medium
2

You are building a Dataflow pipeline in Python that reads messages from Pub/Sub, enriches them with data from a BigQuery table, and writes the results to BigQuery. The enrichment lookup table is large and changes infrequently. Which approach minimizes cost and latency?

Medium
3

A company uses Workflows to orchestrate a series of Google Cloud services for data processing. They need to call an external HTTP API as part of the workflow and handle potential failures with retries. Which Workflows feature should they use?

Medium
4

A media company wants to analyze clickstream data stored in Cloud Storage as JSON files. They need to run SQL queries directly on these files without loading them into BigQuery. Which BigQuery feature should they use?

Easy
5

A team uses dbt on BigQuery to transform data in their data warehouse. They have a large table with nested and repeated fields (arrays and structs). The transformation needs to normalize this data into a star schema. Which dbt feature and BigQuery SQL feature should they use together?

Hard
6

You are building a Dataflow pipeline that reads from Cloud Storage and writes to BigQuery. The pipeline must handle files that are compressed with gzip and contain JSON data. You need to ensure that the pipeline can process these files efficiently and write to BigQuery with minimal errors. Which approach should you take?

Hard
7

Which Dataflow feature automatically scales the number of workers based on the pipeline's current workload, and also selects the optimal machine type for each worker based on the pipeline's resource requirements?

Easy
8

Your team is processing a large dataset with Apache Beam on Dataflow. The pipeline sometimes fails due to transient errors when writing to a BigQuery sink. You need to ensure that failed records are not lost and can be reprocessed later without blocking the pipeline. What is the best approach?

Hard
9

A data engineer needs to load data from CSV files in Cloud Storage into BigQuery. The CSV files have a header row and some columns contain nested JSON strings. Which TWO methods can they use to load this data into BigQuery?

Easy
10

Your organization runs a batch Dataflow pipeline that reads from BigQuery, transforms records, and writes Parquet files to Cloud Storage partitioned by event date. The pipeline currently writes all files into a single directory and downstream Hive-style queries scan the entire dataset. You need to restructure the output so that queries scan only the relevant date partitions, while keeping the pipeline idempotent on re-runs. What should you do?

Medium
11

A data engineer needs to ingest data from a Cloud Storage bucket into BigQuery. The data is in CSV format and is updated daily with new files. The engineer wants to minimize manual intervention and ensure that new files are automatically loaded into BigQuery. Which Google Cloud service should be used to orchestrate this?

Easy
12

An organization needs to continuously replicate change data from a MySQL database to BigQuery with sub-minute latency. The database is running on-premises. Which Google Cloud service should they use?

Medium
13

A data engineer is building a Dataflow pipeline in Python that reads from BigQuery, transforms data, and writes to Cloud Storage. The pipeline will be deployed in production. Which approach should they use to ensure the pipeline is reusable across environments with different configuration parameters?

Medium
14

A data engineer needs to schedule a recurring transfer of data from a partner's Amazon S3 bucket to a Cloud Storage bucket for further processing. Which THREE components or configurations are necessary? (Choose 3)

Easy
15

You are loading 10 GB of daily CSV files from a GCS bucket into a BigQuery table. The files contain some malformed rows that you want to skip. Which BigQuery load configuration should you use?

Easy
16

A data pipeline processes JSON files from Cloud Storage, transforms them using Apache Beam, and writes the output to BigQuery. Some records are malformed and cause the pipeline to fail. How should the engineer handle these errors to ensure the pipeline continues processing while preserving the malformed records for analysis?

Medium
17

You are building a BigQuery table that contains nested and repeated fields (e.g., order with line items). You need to write a query that counts the number of line items per order. Which TWO SQL functions/techniques can you use?

Medium
18

A data engineer needs to migrate 200 TB of on-premises Oracle data to BigQuery. The network bandwidth is limited to 100 Mbps, and the data must be loaded within 2 weeks. Which Google Cloud service is most appropriate for the initial data transfer?

Medium
19

A data engineer needs to transfer 500 TB of archival data from an on-premises NAS to Cloud Storage. The on-premises network has limited bandwidth (100 Mbps). Which transfer method should they recommend?

Easy
20

You need to create a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle malformed messages by writing them to a dead-letter table in BigQuery. Which two Apache Beam transforms or patterns should you use to achieve this? (Choose two.)

Medium
21

A Dataflow batch pipeline reads CSV files from Cloud Storage, joins them with a slowly changing dimension stored in BigQuery, and writes the enriched output to BigQuery. The dimension table is large and the join is causing excessive shuffle and worker memory pressure. The team wants to reduce shuffle while keeping the join logic in the pipeline. Which approach should they use?

Hard
22

You are migrating an on-premises PostgreSQL database to Cloud SQL. You need to continuously replicate changes to BigQuery for real-time analytics with minimal latency. Which service should you use?

Hard
23

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle late-arriving data and emit correct results. You need to ensure that the pipeline's windowing and triggering strategy produces accurate aggregations. Which combination of windowing and triggering should you use?

Hard
24

A company wants to use dbt (data build tool) to transform data in BigQuery. They have a Cloud Storage bucket containing raw CSV files that are loaded daily into BigQuery via an external table. Which dbt feature should they use to modularize the transformation logic and handle dependencies between models?

Medium
25

A healthcare company needs to ingest HL7 messages from an on-premises system into Google Cloud. The messages arrive continuously and must be processed in near-real-time, with transformations applied before loading into BigQuery. The company wants to use a fully managed service for message ingestion and a serverless data processing service. They also need to ensure that the pipeline can handle bursts of traffic and that the data is encrypted at rest. Which TWO Google Cloud services should they use? (Choose two.)

Medium
26

A company wants to move data from an on-premises MySQL database to BigQuery for analytics. They need to capture all changes (inserts, updates, deletes) in near real-time and also perform an initial historical load. Which approach meets these requirements with minimal operational overhead?

Medium
27

You are building a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. You need to ensure that no duplicates are written to BigQuery even if the pipeline retries messages. Which feature should you use?

Easy
28

Which three of the following are valid BigQuery data loading methods? (Choose THREE.)

Medium
29

A company uses Eventarc to trigger a Cloud Run service when new objects appear in a GCS bucket. Recently, the Cloud Run service has been failing with 429 errors (too many requests) during high-velocity uploads. They need to handle the load without losing events. What should they do?

Hard
30

You need to ingest data from a Cloud Storage bucket into BigQuery. The data is in Avro format and you want to minimize the time to insight. Which method should you use?

Easy
31

You are designing a Dataflow pipeline that needs to exactly-once process events from Pub/Sub and write to BigQuery using the Storage Write API. The pipeline may restart and could reprocess some messages. What setting ensures exactly-once semantics for the output?

Hard
32

Your team runs a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. During a deployment, the pipeline is stopped and later restarted with the same job name but without draining. After the restart, downstream reports show duplicate rows for events that were processed just before the stop. You need the pipeline to resume without reprocessing already-published messages. What should you do?

Medium
33

Which TWO statements are true about BigQuery Data Transfer Service? (Choose 2)

Medium
34

A team wants to transfer data from an on-premises Hadoop cluster to Cloud Storage for processing. The cluster is located in a remote area with limited bandwidth. They need to transfer 500 TB of data. Which service should they use?

Medium
35

A company wants to ingest data from an on-premises Oracle database into BigQuery in near real-time with minimal latency. The database has a high volume of inserts and updates. Which service should they use?

Medium
36

A data engineer needs to orchestrate a series of tasks that include calling external APIs, running BigQuery queries, and sending notifications. The workflow involves conditional branching and parallel steps. Which Google Cloud service should be used?

Easy
37

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the Storage Write API with exactly-once semantics. You need to ensure that the pipeline can handle a sudden spike in traffic without losing data or causing duplicates. Which configuration should you adjust to improve throughput while maintaining exactly-once?

Hard
38

You are building a Dataflow pipeline that reads Avro-formatted files from Cloud Storage. The files use a schema that is updated frequently. You want to minimize pipeline restarts due to schema changes. Which approach should you take?

Medium
39

Which BigQuery feature allows you to write data with exactly-once semantics, high throughput, and the ability to buffer data before making it available for queries?

Easy
40

A logistics company uploads nightly shipment manifests as newline-delimited JSON files into a Cloud Storage bucket. Analysts need to query these files with standard SQL immediately after upload, but the schema evolves frequently and the company wants to avoid managing a load job. Which approach should they use?

Medium
41

Your team uses dbt to transform data in BigQuery. You need to schedule dbt runs to refresh materialized tables and views every hour. The transformations include both full refreshes and incremental models. What is the most efficient way to orchestrate these dbt runs on Google Cloud?

Medium
42

A logistics company stores delivery manifests as newline-delimited JSON files in a Cloud Storage bucket. Analysts want to run SQL queries over these files immediately without importing them into a BigQuery dataset or paying for duplicate storage. Which BigQuery capability should they use?

Easy
43

A company needs to load data from a MySQL database into BigQuery daily. The data volume is 10 GB per day and the schema changes occasionally. They want to minimize costs and operational overhead. What is the MOST appropriate approach?

Medium
44

A logistics company streams vehicle telemetry into a Pub/Sub topic. Each message contains a vehicle ID and a timestamp, and the Dataflow pipeline computes per-vehicle distance using a stateful DoFn with a ValueState timer set to fire after 5 minutes of event-time inactivity. During a regional network outage, some vehicles stop sending data for 20 minutes and then resume with correctly ordered timestamps. After the outage, operators notice that some late-arriving records are being dropped before the stateful computation. Which pipeline setting should be adjusted to retain those records for processing?

Hard
45

A data analyst needs to transform nested and repeated fields in BigQuery. They have a table with a column of type ARRAY<STRUCT<...>>. Which SQL function should they use to flatten the array into individual rows for analysis?

Easy
46

You need to process a large volume of event data from Cloud Storage, apply complex transformations using Apache Spark, and then load the results into BigQuery. The data arrives in batches every hour. You want to minimize costs by using preemptible VMs. Which service should you use?

Hard
47

A company uses dbt on BigQuery to transform data. They want to run dbt models on a schedule and manage environments (dev, prod). Which GCP service should they use to run dbt jobs?

Medium
48

Your team is migrating a batch ETL job from an on-premises Hadoop cluster to Dataproc. The job reads CSV files from Cloud Storage, joins them with a slowly changing dimension table in BigQuery, and writes aggregated results back to BigQuery. The on-premises job used Hive on Tez and took six hours. You need to reduce runtime on Dataproc while minimizing cost. Which approach should you take?

Hard
49

A media company stores 400 TB of compressed JSON clickstream logs in a Cloud Storage bucket. Analysts need to run ad hoc SQL over this data several times a day, and the team wants to avoid the cost and delay of loading it into BigQuery. Which BigQuery capability should you use?

Easy
50

You are building a Dataflow pipeline that reads messages from Pub/Sub, performs windowed aggregations, and writes results to BigQuery. The pipeline must handle late-arriving data, but after a certain point, you want to stop accepting late data to allow the pipeline to emit final results. Which Apache Beam concept should you use to define when to stop accepting late data?

Medium
51

You are using Dataproc to run a Spark job that reads data from Cloud Storage, performs aggregations, and writes results back to Cloud Storage. The job is failing with out-of-memory errors on the shuffle. Which optimization should you apply?

Hard
52

A retail analytics team loads point-of-sale transactions into BigQuery every 15 minutes using a load job that appends records. A nightly deduplication job currently removes duplicate transaction IDs, but the team wants duplicates eliminated at ingestion time without adding a separate job. The source system can emit a stable transaction ID for each record. What should you do?

Medium
53

A startup uploads JSON log files to a Cloud Storage bucket throughout the day and wants BigQuery to query them within minutes with no ETL code and no data duplication. The files use newline-delimited JSON and the schema is stable. The team wants the lowest-effort option that still lets queries see newly arrived files automatically. Which approach should they use?

Easy
54

A company wants to trigger a Cloud Run service whenever a new file is uploaded to a specific Cloud Storage bucket. Which event-driven solution should they use?

Easy
55

A media analytics team needs to copy 200 TB of historical log files from an on-premises NFS server into Cloud Storage, then transform them with a Dataflow batch job. The NFS server is reachable only from the corporate network, and the team wants to minimize transfer time and cost. Which approach should they use?

Medium
56

A company wants to transform data using dbt (data build tool) on BigQuery. They have a CI/CD pipeline and need to version-control their transformations. Which setup is recommended?

Medium
57

A data engineer needs to load a 10 GB CSV file from GCS into BigQuery. The file contains some malformed rows that should be skipped. Which approach is most efficient?

Easy
58

A company wants to use Eventarc to trigger a Cloud Run service when new objects are created in a GCS bucket. They also need to filter events for a specific bucket and object prefix. Which THREE resources must exist or be created?

Medium
59

Which Google Cloud service is designed to replicate data from MySQL, PostgreSQL, and Oracle databases to BigQuery or Cloud Storage in near real-time?

Easy
60

Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?

Easy
61

A company wants to orchestrate a multi-step data processing workflow that includes calling a Cloud Run service, waiting for its completion, and then running a BigQuery query. The workflow should be serverless and integrate with Cloud Events. Which Google Cloud service should they use?

Medium
62

A data engineer is designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery using the Storage Write API in exactly-once mode. The pipeline must handle late-arriving data (up to 1 hour) and maintain correct aggregation results. Which trigger configuration should they use?

Hard
63

A company is building a data pipeline that ingests streaming data from Pub/Sub, transforms it with Dataflow, and loads it into BigQuery. They want to handle malformed messages that cannot be parsed. Which TWO actions should they implement for error handling? (Choose 2)

Medium
64

You need to copy a 3 TB dataset from an on-premises Hadoop Distributed File System (HDFS) cluster to Cloud Storage over a 1 Gbps dedicated interconnect. The data is a one-time historical backfill and must be transferred with integrity verification. You want to minimize operational overhead and avoid writing custom transfer code. Which service should you use?

Easy
65

You need to ingest streaming data from a custom application into BigQuery with exactly-once semantics and low latency. The data volume is up to 10 MB/s. Which TWO services should you combine?

Medium
66

A team needs to orchestrate a multi-step workflow that involves calling external APIs, running BigQuery queries, and conditionally executing Cloud Functions. Which Google Cloud service is best suited for this?

Easy
67

A financial firm ingests trade confirmations from a partner's SFTP server into Cloud Storage, then loads them into BigQuery. The partner drops a variable number of files at unpredictable times, and the firm must detect each new file within seconds and trigger a downstream Cloud Run job. The files must not be processed twice if the trigger fires more than once. Which design should you implement?

Hard
68

You need to ingest Google Ads performance data into BigQuery on a daily basis for reporting. Which service should you use?

Easy
69

You need to react to changes in a GCS bucket (e.g., new object creation) and trigger a Cloud Run service to process the new file. Which Google Cloud service should you use to route the event?

Easy
70

A company wants to use dbt to transform data in BigQuery. Their source data is loaded daily into staging tables. They need to run dbt transformations on a schedule and only process tables that have changed. Which dbt feature should they use?

Medium
71

A company wants to stream real-time clickstream data from a website into BigQuery for near-real-time analytics. They expect peaks of 10,000 events per second. Which combination of services is most suitable for ingestion?

Easy
72

A company needs to continuously synchronize customer data changes from an on-premises Oracle database to BigQuery for near-real-time analytics. The Oracle database has Change Data Capture (CDC) enabled. Which Google Cloud service should be used to stream these changes with minimal latency and schema evolution support?

Hard
73

A data engineer needs to ingest on-premises Oracle CDC data into BigQuery in near real-time with minimal operational overhead. Which service should they use?

Easy
74

An organization needs to trigger a Cloud Run service whenever a new file is uploaded to a specific Cloud Storage bucket. Which service should they use to set up this event-driven architecture?

Medium
75

A company is building a real-time anomaly detection pipeline using Dataflow. Events are ingested from Pub/Sub, and the pipeline must compute a sliding window average every minute over a 1-hour window. Which TWO configurations are required for this pipeline? (Choose 2)

Medium
76

A company uses Google Ads and wants to automatically load their advertising data into BigQuery daily. They also need to transform the data with SQL and schedule a recurring query. Which combination of services meets these requirements with minimal operational overhead?

Medium
77

A financial services company receives real-time stock trade data via Pub/Sub. They need to enrich each trade with reference data from a Cloud SQL table and write the results to BigQuery for real-time analytics. The enrichment must handle late-arriving data and ensure exactly-once processing. Which Dataflow streaming pipeline configuration should be used?

Medium
78

A data engineer is building a Dataflow pipeline that reads newline-delimited JSON files from a Cloud Storage bucket and writes them to BigQuery. The files arrive continuously, some with malformed records, and the team wants the pipeline to keep running while routing bad records to a dead-letter location for later analysis. The team also wants the schema to be inferred from the files during development but fixed in production. Which two approaches should the data engineer use? (Choose two.)

Medium
79

A data platform team uses Cloud Data Fusion to move data from an on-premises relational database into BigQuery. They need the pipeline to run on a fixed schedule, capture only rows changed since the last successful run, and avoid re-reading the entire source table each night. The source table has an updated_at column that is reliably populated. Which two approaches should they use? (Choose two.)

Hard
80

A media company streams real-time viewer data from Pub/Sub to BigQuery using a Dataflow pipeline. They need to handle occasional malformed messages without losing valid data. Which pattern should they implement?

Medium
81

You are designing a Dataflow pipeline to process streaming data. The pipeline may encounter malformed records. You need to handle these errors without failing the entire pipeline and store the bad records for later analysis. What is the best practice?

Medium
82

A media analytics team runs a Dataflow streaming pipeline that reads click events from Pub/Sub and writes aggregates to BigQuery. During peak hours, the pipeline's BigQuery write throughput plateaus and Dataflow logs show repeated quota-related retries on the streaming insert API. The team wants to keep exactly-once semantics and increase sustained write throughput. What should they change?

Hard
83

A company wants to migrate 500 TB of on-premises archival data to Cloud Storage. The data is stored on a SAN and the network link is limited to 1 Gbps. The migration must complete within 10 days. What is the MOST cost-effective approach?

Easy
84

A data engineer is using Spark on Dataproc to process a large dataset. They notice the job is slow due to excessive shuffling. They want to optimize the job by using a more efficient data structure that reduces serialization overhead and provides better memory management. Which Spark API should they use?

Hard
85

A company is using Pub/Sub to ingest clickstream events and Dataflow to write to BigQuery. They observe that some events are malformed and cause the pipeline to fail. They need a solution that captures malformed events without blocking the pipeline and allows reprocessing later. Which Dataflow pattern should they implement?

Hard
86

A streaming pipeline ingests events from Pub/Sub, enriches them via a slow REST API call, and writes the result to BigQuery. The API has a limit of 10 requests per second per client. The pipeline processes 1000 messages per second. Which approach minimizes latency while respecting API limits?

Hard
87

A logistics company ingests GPS pings from delivery vans into Pub/Sub, and a Dataflow streaming pipeline writes them to BigQuery. Latency requirements are lenient (about 5 minutes), but the finance team needs the pipeline's cost to be predictable and low, and the data volume fluctuates by a factor of ten between day and night. The team wants to minimize per-element cost without losing data. Which configuration should the data engineer choose?

Hard
88

Your company has a Dataproc cluster that runs Spark jobs. You need to choose between RDDs, DataFrames, and Datasets for a new job that performs complex aggregations on structured data. Which TWO statements are correct regarding performance and ease of use?

Hard
89

A healthcare company needs to process HL7 messages containing sensitive patient data. The messages arrive in Cloud Storage as JSON files. The pipeline must de-identify the data using the Cloud Healthcare API DLP de-identification, then load the results into BigQuery. The pipeline must ensure that no unredacted data is ever written to BigQuery, and that processing is fault-tolerant. The Dataflow pipeline reads from Cloud Storage, calls the DLP API for de-identification, and writes to BigQuery. Which additional configuration ensures that only de-identified data reaches BigQuery?

Hard
90

A media company ingests clickstream events into Pub/Sub and processes them with a Dataflow streaming pipeline that writes to BigQuery. The pipeline uses a fixed window of five minutes and discards late data. A product manager reports that events arriving more than five minutes after their event timestamp never appear in reports. Which change should you make to capture those events?

Hard
91

You need to stream real-time user click events from your application into BigQuery for immediate analysis. The events must be available for query within seconds. Which approach is recommended?

Easy
92

A data engineer needs to query a BigQuery table that contains an array of structs. They want to expand the array into separate rows for each element. Which SQL function should they use?

Easy
93

A data engineer is designing a batch processing pipeline that runs daily. The pipeline reads CSV files from GCS, transforms them using Python, and writes the results to BigQuery. They need to parameterize the pipeline for different environments and run it on a schedule. Which THREE components should they use? (Choose 3)

Medium
94

You are building a Dataflow pipeline that reads from Pub/Sub, applies transformations, and writes to BigQuery. The pipeline must handle late-arriving data and ensure that the windowing and triggering are correct. Which THREE configurations should you consider? (Choose 3)

Hard
95

A company needs to stream real-time user activity data from their application into BigQuery for immediate dashboarding. They want to minimize latency (under 5 seconds) and ensure exactly-once delivery. Which TWO options should they consider? (Choose 2)

Medium
96

A retail company wants to analyze point-of-sale transaction data stored in Cloud SQL for PostgreSQL. They need to run complex analytical queries joining this data with data in BigQuery. The data changes frequently, and they want near-real-time access without impacting the production Cloud SQL instance. Which approach should they use?

Easy
97

A logistics company uploads millions of small JSON files per day to a Cloud Storage bucket and needs to query them with standard SQL immediately, but the analytics team does not want to create BigQuery tables or manage schema changes. The files follow a consistent structure but new fields are added frequently. Which BigQuery capability should they use to query these objects directly from Cloud Storage with the least operational overhead?

Medium
98

A data engineer needs to transfer 500 TB of on-premises data to Google Cloud Storage. The data is stored on NAS devices and the network bandwidth is limited to 100 Mbps. What is the most cost-effective and timely transfer method?

Easy
99

A data engineer is using Apache Spark on Dataproc to process a large dataset. They need to perform complex aggregation and transformation with high performance. The dataset has a known schema and they want to take advantage of Catalyst optimizer. Which Spark API should they use?

Medium
100

A data engineer needs to create a Dataflow pipeline template that can be reused across multiple environments (dev, staging, prod) with different parameters (e.g., input Pub/Sub topic, output BigQuery table). Which template type should they use?

Medium
101

You are designing a streaming pipeline that must handle late-arriving data with a maximum lateness of 10 minutes. You need to ensure that all data is processed exactly once and that results are emitted after the watermark passes the window. Which Apache Beam concept should you use to achieve this?

Hard
102

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. Some incoming messages are malformed and fail to parse. How should you handle these messages to ensure the pipeline continues processing without data loss?

Hard
103

A data engineer needs to schedule recurring nightly loads from Amazon S3 to Google Cloud Storage. The data is in CSV format and the volume is approximately 500 GB per night. Which Google Cloud service should they use?

Easy
104

A media company streams user interaction events into Pub/Sub and processes them with a Dataflow streaming pipeline that writes to BigQuery. During peak hours, the pipeline's watermark lags significantly behind real time, and late-arriving events are being dropped. The team wants late events to be included in windowed aggregations for up to 30 minutes after the window closes. Which Dataflow configuration should they apply?

Hard
105

A financial company uses a Dataflow streaming pipeline to read transactions from Pub/Sub and write to BigQuery. They need exactly-once processing semantics for the BigQuery writes and want to avoid duplicates during pipeline updates. Which approach should they use?

Hard
106

A company uses Workflows to orchestrate a multi-step data pipeline. One step calls an HTTP endpoint that may take up to 10 minutes, but the default Workflows timeout is too short. They also need to handle transient errors with retries. Which TWO configurations should they apply? (Choose 2)

Hard
107

A media analytics company ingests clickstream events into Pub/Sub at a sustained rate of 2 GB/s. A Dataflow streaming pipeline reads these events, performs windowed aggregations, and writes results to BigQuery. The pipeline must handle occasional spikes up to 5 GB/s without data loss or excessive backlog. The operations team wants to minimize manual intervention and cost. What should you do to configure the Dataflow pipeline for dynamic scaling?

Medium

Frequently asked questions

What does the Ingesting and Processing the Data domain cover on the PDE exam?
Be able to design ingestion and processing pipelines that keep running despite bad data, and choose the correct Google Cloud service for each source and trigger. The most important thing: route malformed records to a dead-letter or side output rather than letting them fail the pipeline.
How many questions are in this domain?
This page lists all 107 Ingesting and Processing the Data questions in the PDE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Ingesting and Processing the Data questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pde GOOGLE-PDE pde ingestion processing Practice Questions