Courseiva
Storing the Data →easyMultiple Select

PDE Storing the Data Practice Question

A company wants to implement a data lake on Google Cloud. They need to store raw, structured data in open formats and allow querying directly from BigQuery without loading. Which TWO services or features should they use? (Choose 2)

⚠ Common exam trap

Google often tests the misconception that Dataproc or Dataflow are required for querying data in a data lake, when in fact BigQuery external tables and BigLake provide direct querying without loading, and the key is to recognize that storage (GCS) and the query engine (BigLake) are the correct services.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Storage (GCS)

Cloud Storage (GCS) [CORRECT] is the right choice because a data lake on Google Cloud stores raw, structured data as objects in open formats (e.g., Parquet, Avro, ORC, JSON) in buckets, providing durable, scalable, and cost-effective storage. BigLake [CORRECT] is also correct because it is a storage engine that lets BigQuery query data directly in Cloud Storage (and other sources) without loading it, using external tables with fine-grained access control and metadata caching. Together, GCS holds the raw files and BigLake exposes them to BigQuery for direct querying, which matches the requirement exactly. Dataproc (B) is a managed Spark/Hadoop service for processing, not the storage or direct-query layer, and Dataflow (C) is a managed Apache Beam pipeline service for stream/batch processing, so neither fulfills the storage-plus-query requirement. Cloud SQL (D) is a managed relational database for transactional workloads, not a data lake storage layer, and it does not provide the open-format object storage or BigQuery external querying described.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Cloud Storage (GCS)

    Why this is correct

    GCS is the underlying storage for the data lake, storing data in open formats like Parquet/ORC.

  • ✗

    Dataproc

    Why it's wrong here

    Dataproc is a managed Spark/Hadoop service for processing, not for storing or querying directly.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow is for stream/batch processing, not for direct querying.

  • ✗

    Cloud SQL

    Why it's wrong here

    Cloud SQL is a relational database for OLTP, not a data lake.

  • ✓

    BigLake

    Why this is correct

    BigLake provides a unified interface with fine-grained governance on top of GCS data, enabling external tables in BigQuery.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.