PDE Ingesting and Processing the Data Practice Question
A company needs to stream real-time user activity data from their application into BigQuery for immediate dashboarding. They want to minimize latency (under 5 seconds) and ensure exactly-once delivery. Which TWO options should they consider? (Choose 2)
⚠ Common exam trap
Google Cloud often tests the distinction between legacy streaming inserts (at-least-once) and the Storage Write API in committed mode (exactly-once). Candidates may mistakenly choose legacy inserts because they are simpler to implement, ignoring the exactly-once requirement.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use BigQuery Storage Write API in committed mode
Option B is correct because the BigQuery Storage Write API in committed mode provides exactly-once semantics via write streams and offsets, and it supports low-latency streaming ingestion suitable for sub-5-second dashboarding. Option E is correct because Pub/Sub plus Dataflow is the canonical Google Cloud streaming pipeline: Pub/Sub ingests events durably, and Dataflow's BigQueryIO in STREAMING mode with exactly-once processing guarantees deduplication and exactly-once writes to BigQuery. Option A is not ideal because Cloud Functions invoking the BigQuery REST API (tabledata.insertAll) does not provide exactly-once delivery and adds per-event overhead and cold-start latency. Option C is wrong because legacy streaming inserts offer at-least-once semantics and can produce duplicate rows, violating the exactly-once requirement. Option D is not the best fit because running Kafka on Dataproc adds operational complexity, and the Kafka connector typically relies on the Storage Write API or streaming inserts without inherently guaranteeing exactly-once end-to-end delivery in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud Functions to receive events and call the BigQuery REST API
Why it's wrong here
Invoking the BigQuery REST API from Cloud Functions uses streaming inserts, which provide at-least-once delivery, so duplicate rows remain possible and exactly-once is unmet. It is tempting because Cloud Functions offers simple event-driven ingestion for low-volume workloads where occasional duplicates are tolerable.
- ✓
Use BigQuery Storage Write API in committed mode
Why this is correct
BigQuery Storage Write API in committed mode provides exactly-once semantics through stream-level offsets, so duplicate records are discarded on retry. It supports sub-second, real-time ingestion directly into BigQuery, satisfying the under-five-second latency requirement without intermediate staging. This makes it ideal for streaming live user activity into dashboards.
- ✗
Use BigQuery legacy streaming inserts directly from the application
Why it's wrong here
Legacy streaming inserts offer at-least-once semantics, so duplicate rows can appear and exactly-once delivery is not guaranteed, failing the stated requirement. It is tempting because inserts give low-latency row-level availability, making them suitable when approximate results and immediate dashboarding matter more than deduplication.
- ✗
Use Apache Kafka on Dataproc and write to BigQuery via the BigQuery Kafka connector
Why it's wrong here
Kafka on Dataproc with the BigQuery Kafka connector provides at-least-once delivery by default, so exactly-once is not guaranteed without extra deduplication. It is tempting because Kafka handles high-throughput streaming well, and it is correct when the requirement is durable ordered streaming rather than exactly-once BigQuery writes.
- ✓
Stream data to Pub/Sub, then use Dataflow to write to BigQuery with exactly-once guarantees
Why this is correct
Pub/Sub decouples ingestion from processing, while Dataflow's streaming engine supports exactly-once semantics when writing to BigQuery, satisfying the stem's delivery constraint. Its low-latency pipeline keeps end-to-end delay under five seconds, unlike batch loads, making it suitable for immediate dashboarding of real-time activity.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.