PDE Designing Data Processing Systems Practice Question
A retail company runs a batch pipeline in Cloud Dataflow that reads from Cloud Storage and writes to BigQuery. The pipeline uses a GroupByKey operation on a key that is heavily skewed: one customer ID accounts for 40% of all transactions. This causes a single worker to process a massive amount of data, and the job takes hours longer than expected. The team wants to reduce the skew without changing the pipeline's logic or output. What should they do?
⚠ Common exam trap
The trap here is thinking that adding more workers or larger machines can solve a hot key bottleneck, when the issue is that a single key must be processed by one worker regardless of cluster size.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply a hot key fanout by adding a random secondary key, perform the GroupByKey, then remove the secondary key and re-group.
Hot key fanout is the correct technique for mitigating skew in a GroupByKey. By adding a random secondary key, the hot key's records are spread across multiple workers, and a subsequent grouping removes the secondary key to produce the final result. This preserves the pipeline's semantics while dramatically improving parallelism for the skewed key.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of workers and use a larger machine type for the Dataflow job.
Why it's wrong here
Adding workers or using larger machines does not address the root cause of skew. The hot key still routes all records for that customer to a single worker, so that worker remains the bottleneck. While more workers can handle other keys, the skewed key's processing time is unchanged, and the job will still be delayed by that one key.
- ✗
Replace GroupByKey with CombinePerKey using an associative and commutative CombineFn.
Why it's wrong here
CombinePerKey can reduce data volume before the shuffle, but it does not eliminate skew for a single hot key. The combiner still runs on the worker that receives all records for that key, so the hot key remains a bottleneck. Without adding a secondary key or another splitting mechanism, the skew persists.
- ✓
Apply a hot key fanout by adding a random secondary key, perform the GroupByKey, then remove the secondary key and re-group.
Why this is correct
Hot key fanout splits a single hot key into multiple keys by appending a random suffix, so the GroupByKey distributes the load across many workers. After the first grouping, a second grouping removes the suffix and combines the partial results. This reduces skew without altering the pipeline's logic or output, and it is a standard Dataflow pattern for hot keys.
- ✗
Increase the disk size of the worker VMs to allow more data to be spilled to disk.
Why it's wrong here
Increasing disk size only provides more temporary storage for spilled data; it does not parallelize the processing of the hot key. The single worker still has to read and process all records for that key sequentially. The job may avoid out-of-disk errors, but the processing time remains dominated by the skewed key.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.