mediumMultiple Choice
PDE Practice Question: A team wants to ingest streaming data from…
A team wants to ingest streaming data from millions of IoT devices and store historical data in BigQuery for analysis. They need near real-time analytics on the most recent data, with sub-second latency. Which architecture should they use?
⚠ Common exam trap
Google Cloud often tests the misconception that BigQuery's streaming API can provide sub-second query latency, but in reality, BigQuery is a columnar analytics engine optimized for large scans, not for low-latency point reads, which is why a separate low-latency store like Bigtable is required for real-time access.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Pub/Sub, then a Dataflow pipeline that filters and transforms data, writing to Cloud Bigtable for real-time queries and to Cloud Storage for periodic BigQuery loads.
It uses Cloud Bigtable for sub-second latency on recent data, which is ideal for near real-time analytics on streaming IoT data. Dataflow provides the necessary stream processing, filtering, and transformation before writing to Bigtable for low-latency queries and to Cloud Storage for periodic batch loads into BigQuery for historical analysis. This architecture decouples real-time and historical paths, meeting both latency and storage requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Pub/Sub to receive data, then stream directly into BigQuery using the streaming API, and use standard SQL queries for real-time analytics.
Why it's wrong here
BigQuery's streaming API buffers inserts, so rows are not queryable for seconds, missing sub-second latency. It is tempting because it stores all history in one place with SQL access, and would be correct where near-real-time means seconds rather than sub-second.
- ✓
Use Pub/Sub, then a Dataflow pipeline that filters and transforms data, writing to Cloud Bigtable for real-time queries and to Cloud Storage for periodic BigQuery loads.
Why this is correct
Cloud Bigtable provides single-digit-millisecond row-key lookups, satisfying the sub-second latency requirement for recent data, while Dataflow writes the same stream to Cloud Storage for periodic BigQuery loads that serve historical analytics without straining the real-time path.
- ✗
Use Pub/Sub to ingest data into a Dataproc Spark Streaming job that writes to both Bigtable and BigQuery.
Why it's wrong here
Spark Streaming on Dataproc adds micro-batch scheduling latency, typically seconds, so sub-second reads fail. It is tempting because Bigtable provides fast key lookups alongside BigQuery history, and would suit scenarios needing random-access serving rather than sub-second analytics.
- ✗
Use Cloud SQL to store the latest data and periodically move historical data to BigQuery via cron jobs.
Why it's wrong here
Cloud SQL is a relational OLTP database, not built for millions of device writes per second, and cron-based transfers introduce minutes of lag. It is tempting for small-scale ingestion with periodic reporting, where sub-second latency is not required.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.