Courseiva

PDE Designing Data Processing Systems Practice Question

A data engineer needs to process streaming data from thousands of IoT devices and generate real-time dashboards. The data volume is low but requires exactly-once processing semantics. Which Google Cloud service combination should they use?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Pub/Sub + Cloud Dataflow

Dataflow supports exactly-once processing via its streaming engine and checkpointing. Pub/Sub is the ingest service. Together they provide the required semantics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Pub/Sub + Cloud Data Fusion

    Why it's wrong here

    Cloud Data Fusion is a codeless batch ETL service built on CDAP; it does not natively consume Pub/Sub streams with exactly-once processing for dashboards. It is tempting because it offers graphical pipelines, but that fits batch data integration, not continuous low-volume streaming with exactly-once guarantees.

  • ✗

    Cloud Pub/Sub + Cloud Dataproc

    Why it's wrong here

    Cloud Dataproc runs batch Spark and Hadoop jobs on ephemeral clusters; its streaming support via Spark Streaming does not deliver the managed exactly-once semantics Pub/Sub plus Dataflow provides. It is tempting for existing Spark pipelines, but provisioning clusters suits large batch workloads, not continuous low-volume IoT dashboards.

  • ✓

    Cloud Pub/Sub + Cloud Dataflow

    Why this is correct

    Cloud Pub/Sub ingests the device streams, while Dataflow provides exactly-once processing through its streaming engine's state and deduplication. This satisfies the stem's low-volume, real-time dashboard requirement, since Dataflow's windowing and triggers emit results continuously rather than only on batch completion.

  • ✗

    Cloud Pub/Sub + Cloud Dataprep

    Why it's wrong here

    Cloud Dataprep is a serverless data-preparation tool for interactive cleansing of batch datasets, not a stream processor, so it cannot consume Pub/Sub messages with exactly-once semantics. It is tempting because it integrates with Pub/Sub for ingestion, but that suits exploratory batch wrangling rather than continuous low-volume dashboards.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.