You manage a Cloud Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the BigQueryIO write transform with STREAMING_INSERTS. You need to ensure exactly-once processing semantics for the BigQuery writes. What should you do?
The Storage Write API supports exactly-once semantics when used with Dataflow. Configuring withMethod(STORAGE_WRITE_API) and specifying one or more streams enables the necessary deduplication and transactional guarantees. This is the recommended approach for exactly-once streaming writes to BigQuery, replacing the older STREAMING_INSERTS method.
Why this answer
To achieve exactly-once processing for BigQuery writes in a Dataflow streaming pipeline, use the BigQuery Storage Write API with withMethod(STORAGE_WRITE_API) and configure at least one stream. This method provides transactional writes and deduplication, ensuring each record is written exactly once. Other methods like STREAMING_INSERTS with insertId offer only best-effort deduplication and cannot guarantee exactly-once semantics.
Exam trap
The trap here is assuming that the insertId field in streaming inserts guarantees exactly-once semantics, when in reality it only provides best-effort deduplication within a limited time window.