Courseiva
hardMultiple Choice

PDE Practice Question: Building a data lake on Cloud Storage with data…

A company is building a data lake on Cloud Storage with data from multiple sources. They need to apply schema-on-read and support ad-hoc SQL queries. Which architecture is most suitable?

⚠ Common exam trap

Google Cloud often tests the distinction between schema-on-read (BigQuery external tables) and schema-on-write (traditional databases like Cloud Spanner or Cloud SQL), where candidates mistakenly choose a transactional database for analytical workloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ingest to Cloud Storage, create BigQuery external tables.

BigQuery external tables allow schema-on-read by defining the schema at query time over data stored in Cloud Storage, enabling ad-hoc SQL queries without loading data into a separate system. This architecture directly supports the requirement for schema-on-read and SQL-based analysis, as BigQuery provides a serverless, scalable SQL engine.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Ingest to Cloud Spanner, query directly.

    Why it's wrong here

    Cloud Spanner is a globally distributed OLTP database requiring a fixed schema and does not perform schema-on-read over Cloud Storage objects. It is tempting for its SQL support, but that applies to transactional workloads, whereas the data lake needs external tables or a query engine reading files in situ.

  • ✗

    Ingest to Cloud SQL, then export to Cloud Storage for queries.

    Why it's wrong here

    Cloud SQL is a managed relational database requiring schema-on-write, so ingesting multi-source data there imposes a fixed schema before storage. It is tempting for familiar SQL querying, its actual purpose, but the stem needs schema-on-read against Cloud Storage, which BigQuery external tables deliver.

  • ✓

    Ingest to Cloud Storage, create BigQuery external tables.

    Why this is correct

    BigQuery external tables query Cloud Storage data directly, preserving the raw files without loading or transformation. This satisfies schema-on-read, since the schema is applied at query time rather than ingest, and supports ad-hoc SQL through BigQuery's engine. The multi-source constraint is met because each external table maps to its own source path.

  • ✗

    Ingest to Cloud Storage, load into Dataproc for queries.

    Why it's wrong here

    Loading Cloud Storage data into Dataproc means querying HDFS or cluster-local storage rather than the data lake directly, and it provisions clusters rather than offering serverless ad-hoc SQL. It is tempting for Spark-based batch processing, its real purpose, but BigQuery external tables or Dataproc Serverless query Cloud Storage without loading.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.