easyMultiple Choice
PDE Practice Question: Implement a data lake on Google Cloud to store…
A company wants to implement a data lake on Google Cloud to store raw sensor data (unstructured binary files) and allow data scientists to run SQL queries on processed data. They expect to store terabytes of data and have different access patterns. Which combination of GCP services best meets these requirements?
⚠ Common exam trap
Google Cloud often tests the misconception that Cloud Storage can serve as a queryable database for SQL, when in fact it requires an external query engine like BigQuery or Dataproc for SQL access.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Storage for raw data and BigQuery for processed data
Cloud Storage is the ideal service for storing raw, unstructured binary sensor data at petabyte scale, offering low-cost, durable object storage with multiple access tiers. BigQuery is a serverless, highly scalable data warehouse that allows data scientists to run SQL queries on processed data, with features like columnar storage and automatic optimization for analytical workloads. This combination directly addresses the need for raw storage and SQL-based analytics on processed data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Bigtable for raw data and Cloud Spanner for processed data
Why it's wrong here
Bigtable is a wide-column NoSQL store for low-latency key lookups, not binary object storage, and Spanner is a globally distributed relational database, not an analytics query engine over processed files. This pairing fits high-throughput transactional workloads, not a data lake.
- ✗
Cloud Storage for both raw and processed data
Why it's wrong here
Cloud Storage alone holds the raw binary files but provides no SQL engine, so data scientists cannot query processed data without an additional analytics service. It is the correct landing zone for raw objects, yet the stem also demands SQL querying capability.
- ✗
Cloud SQL for raw data and Cloud Dataproc for processing
Why it's wrong here
Cloud SQL is a managed relational database capped at modest storage, unsuitable for terabytes of raw binary sensor files, and Dataproc is a batch Spark/Hadoop service rather than an interactive SQL warehouse. This suits lift-and-shift relational workloads, not a data lake.
- ✓
Cloud Storage for raw data and BigQuery for processed data
Why this is correct
Cloud Storage holds the raw unstructured binary sensor files cost-effectively at terabyte scale, while BigQuery provides serverless SQL analytics over the processed data. This pairing separates cheap object storage from query compute, matching the differing access patterns described.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.