Courseiva
Fundamental Cloud Concepts →mediumMultiple Choice

Cloud Digital Leader Fundamental Cloud Concepts Practice Question

A data analytics team needs to process streaming data from thousands of IoT devices in real time. They want to ingest the data, process it (e.g., windowed aggregations), and then load it into BigQuery for analysis. Which Google Cloud service should they use for the stream processing step?

⚠ Common exam trap

GCDL often tests the confusion between ingestion (Pub/Sub), processing (Dataflow), and storage/analytics (BigQuery), so candidates may incorrectly select Pub/Sub for stream processing or Dataproc for serverless processing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataflow

Cloud Dataflow is Google Cloud's fully managed, serverless service for both batch and stream processing, built on Apache Beam. It natively supports windowed aggregations, exactly-once processing, and autoscaling, making it ideal for real-time IoT data pipelines. Dataflow can read from Pub/Sub, apply transformations (e.g., fixed/sliding windows), and write to BigQuery with built-in connectors. Unlike Dataproc, it abstracts away cluster management, so the team can focus on pipeline logic rather than infrastructure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery is a serverless data warehouse and analytics engine for SQL queries over vast static datasets. Although it supports a streaming ingestion API for loading data in near real-time, it is not designed to execute complex event-stream transformations such as windowed aggregations, joins across streams, or sessionization. It lacks built-in primitives for stream processing like event-time processing and watermarks, so using it as the primary processing engine for streaming pipelines would require additional orchestration and external logic.

  • ✗

    Cloud Pub/Sub

    Why it's wrong here

    Pub/Sub is a fully managed, asynchronous messaging service that decouples event producers from consumers and provides durable, at-least-once delivery of messages. It acts as the ingestion backbone for streaming data, but it does not compute or transform the data as it flows; it simply buffers and relays messages according to subscription pull/push models. Consequently, relying on Pub/Sub alone to 'process streaming data' would yield no aggregation, filtering, windowing, or analytics — it only moves the raw events.

  • ✗

    Cloud Dataproc

    Why it's wrong here

    Dataproc is a managed service for Apache Hadoop and Spark clusters, which can perform stream processing via Spark Streaming or Structured Streaming using micro-batches. However, that requires you to manage workload scaling, cluster lifecycle, and dependency configuration, and its handling of event-time windows and exactly-once semantics is less integrated than a purpose-built stream processor. For a team needing a seamless, fully managed real-time pipeline with automatic scaling and direct BigQuery integration, Dataproc adds operational overhead and is not as naturally aligned with the event-driven pattern.

  • ✓

    Cloud Dataflow

    Why this is correct

    Dataflow is Google's fully managed, unified programming model for batch and stream processing, built on Apache Beam. It provides exactly-once processing guarantees, intelligent watermarks, and automatic handling of out-of-order data, along with built-in windowing and triggering for time-based aggregations. Its native connectors to BigQuery and Pub/Sub allow the team to read streaming events, transform them in real time, and write results directly into BigQuery for analysis without managing any infrastructure.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

Go deeper

Related to this question

About these practice questions

Courseiva writes every GCDL question from scratch — 848 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.