hardMultiple Choice
PDE Practice Question: Based on the exhibit, what is the most likely…
Exhibit
Refer to the exhibit. ``` # BigQuery table schema and sample data Table: mydataset.events Columns: event_id: STRING (REQUIRED) event_timestamp: TIMESTAMP (REQUIRED) event_data: STRING (NULLABLE) user_id: STRING (REQUIRED) Partitioned by: event_timestamp (daily) Clustered by: user_id Job: Dataflow pipeline writing 1000 events/second to this table using streaming inserts with insertId = event_id. Monitoring shows intermittent 'duplicate rows' in queries that count distinct event_ids. ```
Based on the exhibit, what is the most likely cause of duplicate rows despite using the same event_id as insertId?
⚠ Common exam trap
Google Cloud often tests the misconception that BigQuery's streaming deduplication is a strong guarantee, when in fact it is best-effort and can fail under concurrent writes or short time windows.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BigQuery's streaming buffer deduplication is best-effort and may not catch duplicates within a short time window.
BigQuery's streaming buffer uses best-effort deduplication based on the `insertId` field. When multiple rows are inserted with the same `event_id` mapped to `insertId` within a short time window (typically up to a few minutes), the deduplication mechanism may fail to remove all duplicates, especially under high throughput or network retries. This is a documented limitation of BigQuery streaming, not a guarantee of exactly-once semantics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
BigQuery's streaming buffer deduplication is best-effort and may not catch duplicates within a short time window.
Why this is correct
BigQuery's insertId deduplication operates only within the streaming buffer and is best-effort, not guaranteed. Duplicates arriving in separate buffer windows, or after buffer flush, bypass the check entirely, so identical event_id values still produce duplicate rows.
- ✗
The Dataflow pipeline is retrying inserts due to network errors, and the same event_id is not being used in retries.
Why it's wrong here
BigQuery deduplicates on insertId only within a short window, so retries reusing the same event_id after that window create duplicates; the pipeline must supply a consistent insertId per retry. It is tempting because insertId appears to guarantee exactly-once, which holds only inside the deduplication window.
- ✗
The pipeline is writing more than 100,000 rows per second, exceeding BigQuery's streaming quota.
Why it's wrong here
Streaming quota exhaustion returns RESOURCE_EXHAUSTED errors and forces retries, but BigQuery's insertId deduplication window is best-effort and expires after roughly one minute, so retries beyond that window produce duplicates regardless of throughput. Quota tuning matters when sustained ingestion genuinely exceeds per-second limits.
- ✗
The table is partitioned by timestamp, so BigQuery cannot deduplicate across partitions.
Why it's wrong here
Partitioning does not disable insertId deduplication; BigQuery applies it per table within its best-effort window, and streaming inserts land in the partition derived from the row's timestamp. Partitioning is chosen for query pruning and cost control, not as a deduplication boundary.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.