PDE Ingesting and Processing the Data Practice Question
A Dataflow batch pipeline reads CSV files from Cloud Storage, joins them with a slowly changing dimension stored in BigQuery, and writes the enriched output to BigQuery. The dimension table is large and the join is causing excessive shuffle and worker memory pressure. The team wants to reduce shuffle while keeping the join logic in the pipeline. Which approach should they use?
⚠ Common exam trap
The trap here is reaching for a grouped join such as CoGroupByKey by default, when a side input avoids the shuffle entirely for a broadcastable dimension.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Load the dimension side into a side input and use it in a ParDo to enrich each record.
Using the dimension as a side input lets each worker enrich records locally, avoiding the shuffle and hot-key concentration that CoGroupByKey or a grouped join would introduce. It preserves the in-pipeline join logic and relieves memory pressure. Provisioning more workers or moving the join to BigQuery does not reduce the shuffle as requested.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of workers and raise the disk size per worker to absorb the shuffle.
Why it's wrong here
Adding workers and disk addresses symptoms rather than the shuffle itself, and a large dimension broadcast or grouping can still concentrate load. It increases cost without removing the shuffle bottleneck. The team asked to reduce shuffle, not to provision around it.
- ✗
Use a CoGroupByKey on the two collections and process the grouped results.
Why it's wrong here
CoGroupByKey requires shuffling both collections so records with the same key land on the same worker, which is exactly the shuffle the team wants to avoid. With a large dimension it can also concentrate hot keys on single workers, worsening memory pressure. It does not address the stated problem.
- ✗
Write the CSV data to BigQuery first, then run a SQL join in BigQuery and export the result.
Why it's wrong here
This moves the join out of the pipeline and adds extra load and export steps, contradicting the requirement to keep the join logic in the pipeline. It also introduces additional storage and egress costs. While BigQuery can join efficiently, it is not the requested in-pipeline solution.
- ✓
Load the dimension side into a side input and use it in a ParDo to enrich each record.
Why this is correct
A side input broadcasts the dimension data to every worker so the join happens locally without a shuffle of the main dataset. For a large but manageable dimension table, this eliminates the shuffle that causes memory pressure and speeds up enrichment. It keeps the join logic in the pipeline as required.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.