Courseiva
mediumMultiple Choice

PDE Practice Question: Building a real-time streaming pipeline to ingest…

A company is building a real-time streaming pipeline to ingest clickstream events from web servers, enrich them with user profile data from Cloud Bigtable, and aggregate metrics into BigQuery. The expected throughput is 10,000 events per second with occasional spikes up to 50,000. The data must be processed with low latency (seconds) and exactly-once semantics. Which Google Cloud service should be the core processing engine?

⚠ Common exam trap

Google Cloud often tests the misconception that Cloud Pub/Sub with Cloud Functions is sufficient for low-latency streaming, but candidates overlook that Cloud Functions lacks stateful processing and exactly-once semantics, making it unsuitable for aggregation and enrichment at high throughput.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataflow (Apache Beam runner)

Cloud Dataflow, as a managed Apache Beam runner, is the correct choice because it provides exactly-once processing semantics, low-latency streaming (sub-second to seconds), and autoscaling to handle throughput spikes from 10,000 to 50,000 events per second. Its unified batch and streaming model allows you to enrich clickstream events with user profile data from Cloud Bigtable via side inputs or asynchronous lookups, and write aggregated metrics to BigQuery with exactly-once guarantees using the Beam BigQuery I/O connector.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Cloud Dataflow (Apache Beam runner)

    Why this is correct

    Cloud Dataflow satisfies the low-latency, exactly-once requirement through its Apache Beam runner, which provides native exactly-once processing via streaming state and checkpointing. It autoscales horizontally to absorb the 10,000–50,000 events-per-second spikes, and reads Cloud Bigtable for enrichment before writing aggregates into BigQuery.

  • ✗

    Cloud Pub/Sub with Cloud Functions

    Why it's wrong here

    Cloud Functions is event-driven compute without the stateful windowing, aggregation and exactly-once processing guarantees needed for continuous 50,000 events-per-second streams. It is tempting for lightweight per-event transformation triggered by Pub/Sub, and would suit low-volume, stateless event handling rather than sustained aggregation pipelines.

  • ✗

    Cloud Dataproc with Apache Spark Streaming

    Why it's wrong here

    Cloud Dataproc with Spark Streaming requires provisioning and tuning a cluster, and its micro-batch model complicates exactly-once guarantees under spiky 50,000 events-per-second loads. It is tempting because Spark handles enrichment joins and aggregation, and would suit existing Hadoop or Spark workloads being migrated with minimal code change.

  • ✗

    Cloud Data Fusion

    Why it's wrong here

    Cloud Data Fusion is a managed batch and pipeline orchestration tool built on CDAP, not a low-latency streaming engine delivering exactly-once stateful processing at 50,000 events per second. It is tempting for graphical ETL and enrichment pipelines, and would suit scheduled batch ingestion into BigQuery rather than continuous clickstream processing.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.