PDE Designing Data Processing Systems Practice Question
A startup wants to build a data lake on Google Cloud to store raw JSON, CSV, and Parquet files from various sources. They need a storage solution that is highly durable, globally accessible, and integrates natively with BigQuery and Dataproc. They want to minimize management overhead. Which Google Cloud service should they use?
⚠ Common exam trap
The trap here is assuming that any storage service can serve as a data lake; only object storage like Cloud Storage provides the necessary scale, durability, and native analytics integration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Storage
Cloud Storage is the foundational object store for data lakes on Google Cloud. It offers high durability, global accessibility, and native integration with analytics services like BigQuery and Dataproc. It requires no capacity planning or server management, aligning with the goal of minimal overhead. Relational, NoSQL, and NFS services are not suited for raw file storage at scale.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Filestore
Why it's wrong here
Filestore is a managed NFS file server for applications requiring a shared file system, not an object store for data lake workloads. It is not globally accessible in the same way as Cloud Storage and does not integrate natively with BigQuery external tables. It also requires capacity provisioning, increasing management overhead.
- ✗
Bigtable
Why it's wrong here
Bigtable is a NoSQL wide-column database for high-throughput, low-latency workloads, not a general-purpose file store. It does not natively serve as a data lake for raw files and lacks direct integration with BigQuery external tables. Using it would require data transformation and add complexity.
- ✓
Cloud Storage
Why this is correct
Cloud Storage is a highly durable object store that integrates natively with BigQuery external tables and Dataproc. It requires no provisioning and scales automatically, making it ideal for a data lake with minimal management. It supports all mentioned file formats and is globally accessible.
- ✗
Cloud SQL
Why it's wrong here
Cloud SQL is a managed relational database, not an object store for raw files. It is not designed for storing large JSON, CSV, or Parquet files as a data lake and does not integrate with BigQuery external tables or Dataproc in the same way. It would require schema design and capacity planning, adding management overhead.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.