PDE Ingesting and Processing the Data Practice Question
You need to ingest streaming data from a custom application into BigQuery with exactly-once semantics and low latency. The data volume is up to 10 MB/s. Which TWO services should you combine?
⚠ Common exam trap
PDE often tests the misconception that BigQuery legacy streaming inserts provide exactly-once semantics, when in fact they only offer at-least-once delivery and can duplicate data; candidates must recognize that the Storage Write API is required for exactly-once.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pub/Sub
Option A (Pub/Sub) is correct because it provides a highly scalable, low-latency messaging service that decouples the custom application from downstream processing, ingesting up to 10 MB/s and beyond while buffering the stream reliably. Option D (Dataflow with Storage Write API) is correct because Dataflow can read from Pub/Sub and write to BigQuery using the Storage Write API, which supports exactly-once semantics via its exactly-once delivery mode and provides low-latency, high-throughput ingestion. Together, Pub/Sub plus Dataflow with the Storage Write API form the recommended architecture for exactly-once streaming ingestion into BigQuery. Option B (Cloud Functions) is not appropriate here because it is event-driven and not designed for sustained high-throughput streaming pipelines with exactly-once guarantees. Option C (BigQuery legacy streaming inserts) does not provide exactly-once semantics and is being deprecated in favor of the Storage Write API. Option E (Datastream) is a change data capture (CDC) and replication service for databases, not a general-purpose streaming ingestion path for custom application data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Pub/Sub
Why this is correct
Pub/Sub provides the messaging layer that satisfies the exactly-once delivery constraint: its subscription-level exactly-once delivery feature deduplicates redelivered messages using acknowledgement IDs, so each record reaches BigQuery once. It also sustains the 10 MB/s throughput with low latency, unlike batch-oriented alternatives.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions is an event-driven compute service for running short-lived code in response to triggers; it is not a streaming ingestion pipeline and provides no exactly-once delivery guarantee into BigQuery. It would be correct for lightweight event handling, scheduled tasks or glue logic, not sustained 10 MB/s streaming ingestion.
- ✗
BigQuery legacy streaming inserts
Why it's wrong here
Legacy streaming inserts via the tabledata.insertAll API offer at-least-once delivery, so duplicate rows can appear; exactly-once semantics are not guaranteed. It would be correct for simple, high-throughput append workloads where occasional duplicates are tolerable and best-effort deduplication suffices.
- ✓
Dataflow with Storage Write API
Why this is correct
Dataflow provides exactly-once processing through its streaming engine, while the Storage Write API's exactly-once delivery mode commits rows atomically to BigQuery, eliminating duplicates. Together they satisfy the stem's exactly-once and low-latency constraints at 10 MB/s, which Pub/Sub plus load jobs cannot guarantee.
- ✗
Datastream
Why it's wrong here
Datastream is a serverless change data capture and replication service for databases such as Oracle and MySQL into BigQuery or Cloud Storage; it does not accept arbitrary custom application event streams. It would be correct when replicating existing database changes, not for a bespoke application emitting its own events.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.