mediumMultiple Choice
PDE Practice Question: Building a real-time streaming pipeline to ingest…
A company is building a real-time streaming pipeline to ingest clickstream events from web servers, enrich them with user profile data from Cloud Bigtable, and aggregate metrics into BigQuery. The expected throughput is 10,000 events per second with occasional spikes up to 50,000. The data must be processed with low latency (seconds) and exactly-once semantics. Which Google Cloud service should be the core processing engine?
⚠ Common exam trap
Google Cloud often tests the misconception that Cloud Pub/Sub with Cloud Functions is sufficient for low-latency streaming, but candidates overlook that Cloud Functions lacks stateful processing and exactly-once semantics, making it unsuitable for aggregation and enrichment at high throughput.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataflow (Apache Beam runner)
Cloud Dataflow, as a managed Apache Beam runner, is the correct choice because it provides exactly-once processing semantics, low-latency streaming (sub-second to seconds), and autoscaling to handle throughput spikes from 10,000 to 50,000 events per second. Its unified batch and streaming model allows you to enrich clickstream events with user profile data from Cloud Bigtable via side inputs or asynchronous lookups, and write aggregated metrics to BigQuery with exactly-once guarantees using the Beam BigQuery I/O connector.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Cloud Dataflow (Apache Beam runner)
Why this is correct
Cloud Dataflow satisfies the low-latency, exactly-once requirement through its Apache Beam runner, which provides native exactly-once processing via streaming state and checkpointing. It autoscales horizontally to absorb the 10,000–50,000 events-per-second spikes, and reads Cloud Bigtable for enrichment before writing aggregates into BigQuery.
- ✗
Cloud Pub/Sub with Cloud Functions
Why it's wrong here
Cloud Functions is event-driven compute without the stateful windowing, aggregation and exactly-once processing guarantees needed for continuous 50,000 events-per-second streams. It is tempting for lightweight per-event transformation triggered by Pub/Sub, and would suit low-volume, stateless event handling rather than sustained aggregation pipelines.
- ✗
Cloud Dataproc with Apache Spark Streaming
Why it's wrong here
Cloud Dataproc with Spark Streaming requires provisioning and tuning a cluster, and its micro-batch model complicates exactly-once guarantees under spiky 50,000 events-per-second loads. It is tempting because Spark handles enrichment joins and aggregation, and would suit existing Hadoop or Spark workloads being migrated with minimal code change.
- ✗
Cloud Data Fusion
Why it's wrong here
Cloud Data Fusion is a managed batch and pipeline orchestration tool built on CDAP, not a low-latency streaming engine delivering exactly-once stateful processing at 50,000 events per second. It is tempting for graphical ETL and enrichment pipelines, and would suit scheduled batch ingestion into BigQuery rather than continuous clickstream processing.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.