PDE Designing Data Processing Systems Practice Question
A financial services firm runs a Dataflow batch pipeline that joins a 2 TB transaction dataset with a 40 GB customer reference dataset. The reference data changes only once per day and is currently read from a BigQuery table with a side-input transform on every element. Job cost is dominated by repeated BigQuery reads, and the pipeline occasionally hits quota errors. The team wants to minimize cost and quota pressure while keeping the daily refresh. What should the data engineer change?
⚠ Common exam trap
The trap here is believing that a side input is always cheap, when a large side input is re-materialized per worker and re-read from its source.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Stage the reference dataset as an Avro file in Cloud Storage, load it into the pipeline with a file-based source, and pass it as a side input refreshed by a scheduled daily export from BigQuery.
Moving the slowly changing reference dataset out of BigQuery into a compact Cloud Storage format and refreshing it once daily eliminates read amplification. A single scheduled export replaces thousands of per-element reads, cutting both cost and quota pressure while preserving the daily update cadence. CoGroupByKey, partition pruning on a side input, and BI Engine all leave the repeated BigQuery access pattern intact.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Load the reference dataset into a BigQuery table partitioned by ingestion date and read only the latest partition as the side input.
Why it's wrong here
Partition pruning would reduce bytes scanned, but the side input is still re-read from BigQuery and materialized in worker memory for every bundle or worker that needs it, so repeated API calls and quota consumption persist. The 40 GB dataset is also large for a side input, so this approach risks memory pressure without solving the core read-amplification problem.
- ✗
Enable BigQuery BI Engine on the reference table so side-input queries are served from memory instead of disk.
Why it's wrong here
BI Engine is an in-memory analysis cache designed to accelerate interactive dashboards and sub-second SQL, not to serve Dataflow side-input reads. It has a reservation-based capacity limit and would not eliminate the per-element API calls or quota consumption from the pipeline. It addresses interactive query latency, which is not the bottleneck described in this scenario.
- ✗
Replace the side input with a CoGroupByKey transform that groups transactions and customer records by customer ID before the join.
Why it's wrong here
CoGroupByKey requires both datasets to be keyed PCollections in the same pipeline, so the 40 GB reference table would still have to be read from BigQuery on every run and shuffled. It changes the join mechanics but does not reduce the repeated BigQuery reads or quota consumption that dominate cost here, and it adds a large shuffle that can worsen performance.
- ✓
Stage the reference dataset as an Avro file in Cloud Storage, load it into the pipeline with a file-based source, and pass it as a side input refreshed by a scheduled daily export from BigQuery.
Why this is correct
Exporting the slowly changing 40 GB reference data once per day to Cloud Storage removes per-element BigQuery reads entirely, and Avro gives compact, schema-carrying storage that Dataflow reads efficiently. The single daily export keeps the reference fresh while eliminating the quota pressure and repeated scan cost that the BigQuery side input caused on every pipeline run.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.