PDE Designing Data Processing Systems Practice Question
A data engineer needs to process data in a Dataflow pipeline that reads from a Pub/Sub topic. The pipeline must group events into 5-minute windows and compute the average value per key. Which Beam transform should they use after windowing?
⚠ Common exam trap
The trap is choosing GroupByKey because it 'groups by key,' but the question asks for an aggregation — Combine.perKey is the efficient, purpose-built transform for per-key aggregation and avoids the full shuffle penalty of GroupByKey.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Combine.perKey
Combine.perKey is the correct transform because it performs a per-key aggregation (here, computing the average value) after windowing, combining elements within each key and window efficiently. It is a fused Combine operation that reduces data before shuffling, making it more efficient than GroupByKey followed by a separate aggregation. Since the requirement is to compute an average per key within 5-minute windows, Combine.perKey directly expresses that intent.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Combine.perKey
Why this is correct
Combine.perKey performs per-key aggregation after windowing, computing the average value for each key within each 5-minute window. It satisfies the requirement to group events and average per key, unlike global combines that would merge across keys.
- ✗
ParDo
Why it's wrong here
ParDo applies a user-defined DoFn to each element independently, so it processes events one at a time without collecting all values for a key. Averaging per key requires all elements of that key within the 5-minute window to be gathered first; ParDo alone cannot perform that grouping.
- ✗
GroupByKey
Why it's wrong here
GroupByKey collects all values for each key within the window but emits an iterable of values, not a computed average. A subsequent ParDo with a combining DoFn is still needed to calculate the mean, so GroupByKey alone does not produce the required per-key average.
- ✗
CoGroupByKey
Why it's wrong here
CoGroupByKey performs a relational join across multiple PCollections sharing a window, emitting grouped key-value tuples per input. Computing an average per key needs a single PCollection grouped by key, so CoGroupByKey introduces an unnecessary multi-input join rather than the per-key aggregation the pipeline requires.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.