Courseiva

PDE Preparing and Using Data for Analysis Practice Question

Your company ingests streaming data into BigQuery using the Storage Write API. You need to ensure that duplicate records are not inserted when the streaming job retries due to transient errors. Which feature should you use?

⚠ Common exam trap

Test-takers frequently confuse the legacy streaming API's insertId deduplication with the Storage Write API's exactly-once semantics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the Storage Write API's exactly-once semantics by setting a stream name and offset.

The Storage Write API's exactly-once semantics, enabled by specifying a stream name and offsets, guarantee that records are not duplicated even if the writer retries. This is the only option that provides native deduplication for streaming inserts. Other methods either do not apply to the Storage Write API or are not supported in BigQuery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the Storage Write API's exactly-once semantics by setting a stream name and offset.

    Why this is correct

    The Storage Write API provides exactly-once delivery semantics when you use a stream name and specify offsets for each record. This ensures that retried records are not duplicated. It is designed for high-throughput streaming and supports deduplication natively, making it the correct choice for preventing duplicates on retries.

  • ✗

    Create a unique constraint on the destination table to reject duplicate rows.

    Why it's wrong here

    BigQuery does not support unique constraints or primary keys. While you can enforce uniqueness via application logic or deduplication queries, there is no native constraint that rejects duplicates on insert. This approach is not feasible in BigQuery and would not prevent duplicates during streaming.

  • ✗

    Enable BigQuery's streaming inserts with insertId to deduplicate records.

    Why it's wrong here

    The legacy streaming API uses insertId for best-effort deduplication, but it is not guaranteed and is not supported in the Storage Write API. The Storage Write API uses a different mechanism for exactly-once semantics. Relying on insertId would not meet the requirement for guaranteed deduplication in this scenario.

  • ✗

    Use a MERGE statement to upsert records after each streaming batch.

    Why it's wrong here

    MERGE statements are used for batch upserts and are not suitable for streaming inserts because they require running a query over existing data, adding latency and cost. They also do not prevent duplicates during the streaming write itself. For streaming deduplication, the Storage Write API's exactly-once feature is more appropriate.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.