PDE Designing Data Processing Systems Practice Question
A data engineer needs to process streaming data from thousands of IoT devices and generate real-time dashboards. The data volume is low but requires exactly-once processing semantics. Which Google Cloud service combination should they use?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Pub/Sub + Cloud Dataflow
Dataflow supports exactly-once processing via its streaming engine and checkpointing. Pub/Sub is the ingest service. Together they provide the required semantics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Pub/Sub + Cloud Data Fusion
Why it's wrong here
Cloud Data Fusion is a codeless batch ETL service built on CDAP; it does not natively consume Pub/Sub streams with exactly-once processing for dashboards. It is tempting because it offers graphical pipelines, but that fits batch data integration, not continuous low-volume streaming with exactly-once guarantees.
- ✗
Cloud Pub/Sub + Cloud Dataproc
Why it's wrong here
Cloud Dataproc runs batch Spark and Hadoop jobs on ephemeral clusters; its streaming support via Spark Streaming does not deliver the managed exactly-once semantics Pub/Sub plus Dataflow provides. It is tempting for existing Spark pipelines, but provisioning clusters suits large batch workloads, not continuous low-volume IoT dashboards.
- ✓
Cloud Pub/Sub + Cloud Dataflow
Why this is correct
Cloud Pub/Sub ingests the device streams, while Dataflow provides exactly-once processing through its streaming engine's state and deduplication. This satisfies the stem's low-volume, real-time dashboard requirement, since Dataflow's windowing and triggers emit results continuously rather than only on batch completion.
- ✗
Cloud Pub/Sub + Cloud Dataprep
Why it's wrong here
Cloud Dataprep is a serverless data-preparation tool for interactive cleansing of batch datasets, not a stream processor, so it cannot consume Pub/Sub messages with exactly-once semantics. It is tempting because it integrates with Pub/Sub for ingestion, but that suits exploratory batch wrangling rather than continuous low-volume dashboards.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.