Courseiva

PDE · topic practice

Ingesting and Processing the Data practice questions

This domain covers moving data into and through Google Cloud: Pub/Sub ingestion, Dataflow/Apache Beam pipelines, Cloud Storage event triggers, BigQuery loading, Dataproc, and transfer options like Storage Transfer Service and Transfer Appliance. Questions are scenario-based, asking you to pick the right service, handle failures such as malformed records, and design pipelines that keep running rather than aborting.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Ingesting and Processing the Data

What the exam tests

What to know about Ingesting and Processing the Data

Be able to design ingestion and processing pipelines that keep running despite bad data, and choose the correct Google Cloud service for each source and trigger. The most important thing: route malformed records to a dead-letter or side output rather than letting them fail the pipeline.

Using Dataflow dead-letter patterns and side outputs to route malformed Pub/Sub or JSON records without failing the pipeline

Choosing Pub/Sub, Eventarc, or Cloud Storage notifications to trigger Cloud Run or Cloud Functions on object events

Selecting Storage Transfer Service versus Transfer Appliance for large on-premises Hadoop-to-Cloud Storage migrations

Building Apache Beam pipelines that read Cloud Storage, transform, and write to BigQuery with correct windowing and error handling

Watch out for

Common Ingesting and Processing the Data exam traps

  • ▸Letting a single malformed record throw an exception that kills the whole Dataflow pipeline instead of routing bad records to a dead-letter sink or side output.
  • ▸Assuming Pub/Sub or Cloud Storage triggers automatically invoke Cloud Run without configuring Eventarc or notifications and the required IAM permissions.
  • ▸Picking Transfer Appliance for a transfer that fits within network bandwidth and deadline, when Storage Transfer Service over the network would suffice.

Practice set

Ingesting and Processing the Data questions

20 questions · select your answer, then reveal the explanation

A data engineer needs to load 10 TB of CSV files from Amazon S3 into Google BigQuery on a daily basis. Which service should they use to automate this transfer?

Your company uses Kafka for event streaming. You want to run Kafka on Google Cloud with the ability to auto-scale clusters and use managed infrastructure. Which service should you choose?

You need to perform a one-time migration of historical data from an on-premises Teradata data warehouse to BigQuery. The data volume is 50 TB and you have a high-speed network connection (10 Gbps). What is the most efficient way to load the data?

You have a Dataflow pipeline that processes streaming data with high throughput. You notice that the pipeline is experiencing high latency and the workers are underutilized. Which Dataflow feature can automatically optimize resource allocation?

Your organization uses dbt (data build tool) for transformations on BigQuery. You need to run dbt models on a schedule and manage versions. Which Google Cloud service can execute dbt jobs in a serverless manner?

An organization wants to ingest on-premises Oracle database changes into BigQuery for real-time analytics with minimal latency. The Oracle database is version 19c and has a high transactional volume. Which Google Cloud service should they use?

A company runs a Dataflow pipeline that reads from Pub/Sub, transforms data, and writes to BigQuery. The pipeline uses classic templates and is deployed in batch mode. They notice that the pipeline does not scale well under high load, causing a backlog in Pub/Sub. Which improvement would BEST address the scaling issue?

An organization needs to transfer 50 TB of historical data from an on-premises Hadoop cluster to Google Cloud Storage. The network bandwidth is limited to 100 Mbps. Which transfer method is MOST cost-effective and time-efficient?

A data engineer is designing a Dataflow pipeline in Python that reads from Pub/Sub, applies complex transformations using external libraries, and writes to BigQuery. The pipeline must be deployed as a reusable, version-controlled template that can be easily updated without re-uploading the pipeline code each time. Which approach should they use?

A financial services company needs to ingest real-time trade data from multiple sources into BigQuery for immediate fraud detection. The data volume is high (1 million messages per second) and each message must be available for queries within seconds. They are considering the Storage Write API. Which stream mode should they choose to balance data availability and cost?

A data engineering team needs to ingest streaming data from an existing Kafka cluster (on-premises) into Google Cloud for real-time analytics. They want to minimize changes to the existing Kafka setup and avoid long-term operational overhead. Which TWO approaches should they consider?

A large enterprise is migrating its data warehouse from Teradata to BigQuery. They need to transfer historical data (100 TB) and set up ongoing daily incremental loads. They also need to transform the data using dbt. Which THREE Google Cloud services should they use?

You are building a streaming pipeline to ingest real-time clickstream data from a website into BigQuery for immediate analysis. The data must be available in BigQuery within seconds and you need to handle late-arriving data (e.g., browser offline events) that may arrive hours later. Which approach should you use?

A company wants to transfer 500 TB of data from an on-premises Hadoop cluster to Google Cloud Storage (GCS) for processing with Dataproc. The on-premises network has a 1 Gbps dedicated link to Google Cloud. The data must be transferred as quickly as possible, minimizing network usage. Which transfer method should they use?

You are designing a near-real-time CDC pipeline to replicate changes from an on-premises PostgreSQL database to BigQuery for analytics. The source database has high transaction volume and you must ensure minimal impact on the source. Which Google Cloud service should you use to ingest the change data?

Your Dataflow pipeline reads from Pub/Sub, performs transformations, and writes to BigQuery. You notice that the pipeline's autoscaling is not keeping up with sudden spikes in traffic, causing increased lag. The pipeline uses Classic Templates. Which change would most effectively improve autoscaling responsiveness?

You are migrating an existing Kafka cluster to Google Cloud using Dataproc. The cluster handles high-throughput streaming data with strict ordering requirements per partition. Which choice of Dataproc configuration is most appropriate?

You are designing a streaming pipeline that ingests events from Pub/Sub, enriches them with a machine learning model, and writes the results to BigQuery. The ML model is deployed on Cloud Run and has a high latency (500ms per request). You need to minimize the impact of slow ML inference on the overall pipeline throughput. Which approach should you take?

Your team has a Dataflow pipeline that reads from BigQuery, transforms data, and writes to GCS. The pipeline is failing with 'Out of Memory' errors on the worker nodes. The input data is large but fits within the total cluster memory. Which configuration change is most likely to resolve the issue without increasing costs significantly?

A company wants to stream real-time user click events from their web application into BigQuery for immediate analysis. Which combination of services is the most scalable and cost-effective for this use case?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Ingesting and Processing the Data sessions

Start a Ingesting and Processing the Data only practice session

Every question in these sessions is drawn from the Ingesting and Processing the Data domain — nothing else.

Related practice questions

Related PDE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PDE exam test about Ingesting and Processing the Data?
Be able to design ingestion and processing pipelines that keep running despite bad data, and choose the correct Google Cloud service for each source and trigger. The most important thing: route malformed records to a dead-letter or side output rather than letting them fail the pipeline.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Ingesting and Processing the Data questions in a focused session?
Yes — the session launcher on this page draws every question from the Ingesting and Processing the Data domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PDE topics?
Use the topic links above to move to related areas, or go back to the PDE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PDE exam covers. They are not copied from any real exam or dump site.