PDE Storing the Data Practice Question
A company wants to implement a data lake on Google Cloud. They need to store raw, structured data in open formats and allow querying directly from BigQuery without loading. Which TWO services or features should they use? (Choose 2)
⚠ Common exam trap
Google often tests the misconception that Dataproc or Dataflow are required for querying data in a data lake, when in fact BigQuery external tables and BigLake provide direct querying without loading, and the key is to recognize that storage (GCS) and the query engine (BigLake) are the correct services.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Storage (GCS)
Cloud Storage (GCS) [CORRECT] is the right choice because a data lake on Google Cloud stores raw, structured data as objects in open formats (e.g., Parquet, Avro, ORC, JSON) in buckets, providing durable, scalable, and cost-effective storage. BigLake [CORRECT] is also correct because it is a storage engine that lets BigQuery query data directly in Cloud Storage (and other sources) without loading it, using external tables with fine-grained access control and metadata caching. Together, GCS holds the raw files and BigLake exposes them to BigQuery for direct querying, which matches the requirement exactly. Dataproc (B) is a managed Spark/Hadoop service for processing, not the storage or direct-query layer, and Dataflow (C) is a managed Apache Beam pipeline service for stream/batch processing, so neither fulfills the storage-plus-query requirement. Cloud SQL (D) is a managed relational database for transactional workloads, not a data lake storage layer, and it does not provide the open-format object storage or BigQuery external querying described.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Cloud Storage (GCS)
Why this is correct
GCS is the underlying storage for the data lake, storing data in open formats like Parquet/ORC.
- ✗
Dataproc
Why it's wrong here
Dataproc is a managed Spark/Hadoop service for processing, not for storing or querying directly.
- ✗
Dataflow
Why it's wrong here
Dataflow is for stream/batch processing, not for direct querying.
- ✗
Cloud SQL
Why it's wrong here
Cloud SQL is a relational database for OLTP, not a data lake.
- ✓
BigLake
Why this is correct
BigLake provides a unified interface with fine-grained governance on top of GCS data, enabling external tables in BigQuery.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.