PDE Ingesting and Processing the Data Practice Question
A company wants to stream real-time clickstream data from a website into BigQuery for near-real-time analytics. They expect peaks of 10,000 events per second. Which combination of services is most suitable for ingestion?
⚠ Common exam trap
The trap is assuming that any streaming combination works equally well, but the exam expects knowledge that the Storage Write API is preferred over legacy streaming inserts for performance and exactly-once semantics, and that Pub/Sub is essential for handling peak loads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pub/Sub → Dataflow → BigQuery (Storage Write API)
For high-throughput, real-time streaming into BigQuery, the combination of Pub/Sub for ingestion, Dataflow for stream processing, and BigQuery's Storage Write API for writing is the most suitable. Pub/Sub handles the 10,000 events per second peaks reliably, Dataflow provides scalable stream processing, and the Storage Write API offers exactly-once semantics and better performance than legacy streaming inserts.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Storage → Cloud Functions → BigQuery
Why it's wrong here
Cloud Functions invoked per object gives batch, event-driven processing, not continuous streaming, and Cloud Storage adds latency before BigQuery sees rows. It tempts for low-volume file drops, where a function triggered on upload is a legitimate ingestion pattern.
- ✗
Direct Web → Dataflow → BigQuery
Why it's wrong here
Dataflow cannot ingest HTTP requests directly; it needs a source such as Pub/Sub, so the website has no endpoint to publish to. It tempts because Dataflow handles windowing and exactly-once processing downstream, which suits this volume once events reach a queue.
- ✓
Pub/Sub → Dataflow → BigQuery (Storage Write API)
Why this is correct
Pub/Sub decouples ingestion from processing, absorbing 10,000 events per second spikes without loss, while Dataflow provides exactly-once streaming transformation. The Storage Write API commits rows directly into BigQuery, satisfying the near-real-time analytics requirement without staging files in Cloud Storage.
- ✗
Pub/Sub → Dataflow → BigQuery (legacy streaming inserts)
Why it's wrong here
Legacy streaming inserts via tabledata.insertAll lack the exactly-once semantics of the Storage Write API, risking duplicates at 10,000 events per second. It tempts because Pub/Sub plus Dataflow is otherwise the canonical pipeline; only the BigQuery write method is outdated.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.