Courseiva
mediumMultiple Choice

PDE Practice Question: Building a real-time streaming pipeline using…

A company is building a real-time streaming pipeline using Pub/Sub and Dataflow to process clickstream data. The pipeline writes aggregated metrics to BigQuery every 10 seconds using a fixed window. During peak traffic, some windows produce duplicate rows in BigQuery. What is the most likely cause?

⚠ Common exam trap

Test-takers frequently confuse trigger behavior (Option B) with the root cause of duplicates, not realizing that duplicates stem from retry semantics in the sink, not from windowing or parallelism.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataflow is retrying BigQuery streaming inserts after a timeout, and the retries succeed even though the original insert succeeded.

Dataflow uses at-least-once semantics for streaming inserts into BigQuery. When a streaming insert times out, Dataflow retries the insert, and if the original insert actually succeeded but the acknowledgment was lost, the retry produces a duplicate row. This is a known behavior of BigQuery streaming inserts with retry logic.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Dataflow is retrying BigQuery streaming inserts after a timeout, and the retries succeed even though the original insert succeeded.

    Why this is correct

    BigQuery streaming inserts are at-least-once, not exactly-once. When Dataflow's insert call times out, the bundle is retried; if the original insert actually committed, the retry writes the same window's rows again, producing duplicates. Idempotent deduplication keys or BigQuery Storage Write API with exactly-once semantics prevent this.

  • ✗

    The pipeline uses default triggers instead of after-watermark triggers.

    Why it's wrong here

    Default triggers fire repeatedly on early and late panes, emitting multiple results per window; after-watermark triggers emit once when the watermark passes, preventing duplicates. Default triggering suits speculative low-latency output where downstream deduplication is acceptable, not exactly-once BigQuery writes.

  • ✗

    The fixed window duration is too short, causing overlapping windows.

    Why it's wrong here

    Fixed windows of equal length never overlap; each element belongs to exactly one window, so shortening the duration cannot create duplicates. Duplicates arise from repeated trigger firings per window. Short windows suit low-latency dashboards, but they change emission frequency, not row duplication.

  • ✗

    The pipeline is using too many Dataflow workers, causing load balancing issues.

    Why it's wrong here

    Worker count affects parallelism and throughput, not result correctness; Dataflow's shuffle and checkpointing keep per-key state consistent regardless of worker scaling. Duplicate rows stem from trigger behaviour emitting multiple panes. Adding workers suits backlog reduction during peak load, not fixing duplicate output.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.