Courseiva
mediumMultiple SelectObjective-mapped

PDE Designing a data lake on Google Cloud Practice Question

A company is designing a data lake on Google Cloud. They need to store raw data in multiple formats (CSV, Parquet, Avro) and allow various downstream processing frameworks. Which two storage solutions provide flexibility and scalability? (Choose two.)

⚠ Common exam trap

Google Cloud often tests the misconception that any database or storage service can serve as a data lake, but the trap here is that only object storage (Cloud Storage) and a serverless query engine (BigQuery) provide the schema-on-read flexibility and scalability required for raw multi-format data, while transactional or operational databases (Spanner, Bigtable) impose schema-on-write constraints and are not designed for bulk analytical storage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

BigQuery

BigQuery is correct because it can directly query raw data stored in Cloud Storage in formats like CSV, Parquet, and Avro using external tables or federated queries, without requiring data loading. This provides a flexible, serverless analytics layer that scales automatically and integrates with downstream processing frameworks like Apache Spark, Dataflow, and Dataproc.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Cloud Filestore

    Why it's wrong here

    Filestore provides shared filesystems for VMs, not scalable object storage.

  • BigQuery

    Why this is correct

    BigQuery can store and query structured data, and with federated queries it can access external files.

  • Cloud Storage

    Why this is correct

    Cloud Storage is object storage that can hold any file format and is highly scalable.

  • Cloud Spanner

    Why it's wrong here

    Spanner is a relational database with strong consistency, not for data lake storage.

  • Cloud Bigtable

    Why it's wrong here

    Bigtable is a NoSQL wide-column database, not suitable for raw file storage.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 890 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.