easyMultiple Choice
End-to-End Exactly-Once Processing — Pub/Sub to Bigtable
A company is ingesting real-time sensor data from thousands of devices into Cloud Pub/Sub. They need to process this data with low latency (seconds) and exactly-once semantics. Which data processing service should they use?
⚠ Common exam trap
Google Cloud often tests the misconception that serverless services like Cloud Functions or Cloud Run inherently provide exactly-once processing, when in fact they rely on Pub/Sub's at-least-once delivery and require additional logic to achieve exactly-once semantics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataflow streaming with exactly-once processing
Dataflow streaming with exactly-once processing is the correct choice because it provides exactly-once semantics for Pub/Sub sources via checkpointing and idempotent sinks, and it meets the low-latency (seconds) requirement through its streaming engine that minimizes per-element overhead. Cloud Dataflow's integration with Pub/Sub ensures that each message is processed exactly once, even in the presence of failures, by using snapshots and consistent state management.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Run with Pub/Sub push
Why it's wrong here
Cloud Run with Pub/Sub push delivers at-least-once; duplicate deliveries occur, so exactly-once semantics are not provided. It suits stateless HTTP or event handlers where occasional reprocessing is tolerable, not stateful streaming with deduplication and windowing requirements.
- ✗
Cloud Functions triggered by Pub/Sub
Why it's wrong here
Cloud Functions triggered by Pub/Sub also delivers at-least-once, so duplicate invocations break exactly-once processing. It fits lightweight, short-lived event handlers, whereas exactly-once streaming aggregation needs a runner that tracks state and checkpoints offsets, such as Dataflow.
- ✓
Dataflow streaming with exactly-once processing
Why this is correct
Dataflow streaming provides the low-latency, per-record processing Pub/Sub ingestion requires, and its exactly-once mode deduplicates via Pub/Sub message IDs and transactional state writes. This satisfies both the seconds-level latency and exactly-once semantics constraints without custom checkpointing.
- ✗
Dataproc with Spark Streaming
Why it's wrong here
Spark Streaming on Dataproc provides at-least-once processing by default; exactly-once requires extra idempotent sinks or transactional writes you must engineer yourself. Dataproc suits lift-and-shift Hadoop or Spark workloads, not a managed service offering exactly-once semantics out of the box.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.