PDE Preparing and Using Data for Analysis Practice Question
Your team uses Cloud Dataflow to stream events from Pub/Sub into BigQuery. Some events arrive late, up to 10 minutes after their event timestamp. You need the pipeline to produce correct aggregations per 5-minute window, including late data, while keeping latency low for on-time events. What should you do?
⚠ Common exam trap
Candidates often confuse Pub/Sub acknowledgment deadlines or sink exactly-once settings with Dataflow windowing semantics, which are separate concerns.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the windowing strategy to fixed windows of 5 minutes and configure allowed lateness to 10 minutes with a trigger that emits early results and updates them on late arrivals.
To correctly aggregate late events in Dataflow, you must use event-time windows with allowed lateness so state is retained for late arrivals, and triggers to emit early and updated results. Sink configuration and subscription settings do not affect windowing semantics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the pipeline's default expansion service to use the BigQuery Storage Write API and enable exactly-once semantics for the sink.
Why it's wrong here
The sink configuration controls how data is written to BigQuery, not how windows handle late events. Exactly-once semantics prevent duplicate writes but do not correct aggregations for late-arriving data. The windowing and trigger configuration is what determines whether late events are included.
- ✓
Set the windowing strategy to fixed windows of 5 minutes and configure allowed lateness to 10 minutes with a trigger that emits early results and updates them on late arrivals.
Why this is correct
Allowed lateness tells Dataflow to keep window state for 10 minutes after the watermark passes, so late events are included. Using early triggers provides low-latency partial results, and late triggers update them. This satisfies both correctness for late data and low latency for on-time events.
- ✗
Configure the Pub/Sub subscription to have a 10-minute acknowledgment deadline and increase the Dataflow worker count to handle the backlog.
Why it's wrong here
The acknowledgment deadline affects how long Pub/Sub waits for ack, not how Dataflow handles late-arriving event timestamps. Increasing workers improves throughput but does not change windowing semantics. This option does not address the need to include late data in per-window aggregations.
- ✗
Use a global window with a repeating trigger every 5 minutes and set the accumulation mode to DISCARDING.
Why it's wrong here
A global window does not group events by their event-time 5-minute intervals, so aggregations would not be per-window as required. DISCARDING mode also drops previous results, losing late data corrections. This approach does not produce correct per-window aggregations for late events.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.