Courseiva

PDE Preparing and Using Data for Analysis Practice Question

A data engineer needs to design a data pipeline that ingests streaming data from Cloud Pub/Sub, performs real-time aggregations, and loads the results into BigQuery for dashboarding. Which Google Cloud service should they use for the streaming aggregation step?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataflow

Dataflow is a fully managed service for stream and batch processing that integrates with Pub/Sub and BigQuery. It supports exactly-once processing and low-latency streaming.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Functions

    Why it's wrong here

    Cloud Functions is event-driven and stateless, so it cannot hold the windowed state that real-time aggregation across streaming records requires. It suits short, single-event triggers such as reacting to a file upload, not continuous windowed computation.

  • ✓

    Dataflow

    Why this is correct

    Dataflow provides serverless, autoscaling Apache Beam processing with exactly-once semantics, satisfying the stem's real-time aggregation requirement. Its native Pub/Sub source and BigQuery sink connectors handle windowed streaming aggregations without managing compute infrastructure, unlike Dataproc's cluster-based Spark or BigQuery's load-only ingestion.

  • ✗

    Cloud Dataproc

    Why it's wrong here

    Dataproc runs batch Spark and Hadoop jobs on provisioned clusters, not continuous Pub/Sub windowed aggregation; its micro-batch latency and cluster lifecycle suit migration or periodic ETL workloads, whereas Dataflow provides native streaming windows and exactly-once BigQuery sinks.

  • ✗

    Cloud Composer

    Why it's wrong here

    Composer orchestrates workflows via Airflow DAGs; it schedules and triggers jobs but does not itself process Pub/Sub messages or compute windowed aggregations. It would be correct for coordinating dependent batch pipelines, not for the real-time transformation stage.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.