PDE Designing Data Processing Systems Practice Question
You are designing a data pipeline that processes streaming events with late-arriving data (up to 2 hours late). The pipeline must compute hourly aggregations and emit results as soon as possible, but must also accurately update results when late data arrives. You want to minimize overall processing cost. Which Dataflow windowing and trigger configuration should you use?
⚠ Common exam trap
The trap is misunderstanding triggers and allowed lateness; candidates may choose global windows or sliding windows without considering the need for hourly aggregations and late data handling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Fixed windows of 1 hour with allowed lateness of 2 hours and trigger every 5 minutes (early) and on watermark (late) with accumulating fired panes
Fixed windows of 1 hour with allowed lateness of 2 hours and triggers every 5 minutes (early) and on watermark (late) with accumulating fired panes is the correct configuration. This setup computes hourly aggregations, emits early results every 5 minutes, and updates results when late data arrives within the 2-hour allowed lateness. Accumulating panes ensure that late data updates the previous results. This minimizes cost by using fixed windows and appropriate triggers.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Fixed windows of 1 hour with allowed lateness of 2 hours and trigger every 5 minutes (early) and on watermark (late) with accumulating fired panes
Why this is correct
Fixed one-hour windows with two hours of allowed lateness retain late events, while early periodic triggers plus a watermark trigger emit results promptly and accumulating panes revise prior output. This satisfies both low-latency emission and accurate late-data correction without over-provisioning resources.
- ✗
Global window with triggers every 5 minutes
Why it's wrong here
A global window with 5-minute triggers cannot produce hourly aggregations, and it never closes, so late events arriving up to two hours later cannot be reconciled into completed hourly results. It suits continuous low-latency monitoring of unbounded streams, not per-hour emission with accurate late-data updates.
- ✗
Sliding windows of 1 hour with 30-minute offset
Why it's wrong here
Sliding windows recompute overlapping aggregates continuously, multiplying state and emission cost, and a fixed 30-minute offset cannot accommodate arrivals up to two hours late. Sliding windows suit rolling metrics like moving averages, not hourly aggregation with late-data corrections.
- ✗
Session windows with 10-minute gap duration
Why it's wrong here
Session windows key on inactivity gaps, so a 10-minute gap fragments a continuous event stream into many small windows instead of hourly aggregates, and late arrivals cannot reliably reopen them. Session windows suit bursty, user-activity grouping, not fixed hourly aggregation.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.