Courseiva

PDE Ingesting and Processing the Data Practice Question

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the Storage Write API with exactly-once semantics. You need to ensure that the pipeline can handle a sudden spike in traffic without losing data or causing duplicates. Which configuration should you adjust to improve throughput while maintaining exactly-once?

⚠ Common exam trap

The trap here is assuming that scaling the Dataflow workers alone will increase BigQuery write throughput, but the Storage Write API has its own stream-based throughput limit that must be tuned separately.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Increase the number of Storage Write API streams.

The Storage Write API uses a configurable number of streams to write data to BigQuery. Each stream can handle a certain throughput, and increasing the number of streams allows more parallel writes, improving overall throughput. This adjustment maintains exactly-once semantics because each stream is managed with its own offset and deduplication. Autoscaling and worker count help with processing but do not directly increase the write throughput limit imposed by the number of streams.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Increase the number of Storage Write API streams.

    Why this is correct

    The Storage Write API allows multiple streams to be used concurrently. Increasing the number of streams (via withNumStorageWriteApiStreams) can improve throughput by parallelizing writes. This is a key tuning parameter for high-volume pipelines while preserving exactly-once semantics, as each stream maintains its own offset and deduplication.

  • ✗

    Increase the maximum number of workers.

    Why it's wrong here

    Increasing the maximum number of workers can provide more compute resources for parallel processing, but if the Storage Write API is configured with a fixed number of streams, the write throughput may not scale. The number of streams is a separate limit that must be raised to fully utilize additional workers.

  • ✗

    Enable autoscaling for the Dataflow job.

    Why it's wrong here

    Autoscaling adjusts the number of worker instances based on load, which can help with CPU-bound processing but does not directly address the write throughput limit of the Storage Write API. Without increasing the number of streams, the write stage may still be a bottleneck. Autoscaling alone may not suffice for a sudden spike.

  • ✗

    Switch to STREAMING_INSERTS method.

    Why it's wrong here

    Switching to STREAMING_INSERTS would sacrifice exactly-once semantics because that method provides at-least-once delivery and can introduce duplicates. It may offer higher throughput in some cases, but it does not meet the exactly-once requirement. The Storage Write API is the correct method, and tuning streams is the proper approach.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.