Courseiva
easyMultiple Choice

PDE Practice Question: Implement a data lake on Google Cloud to store…

A company wants to implement a data lake on Google Cloud to store raw sensor data (unstructured binary files) and allow data scientists to run SQL queries on processed data. They expect to store terabytes of data and have different access patterns. Which combination of GCP services best meets these requirements?

⚠ Common exam trap

Google Cloud often tests the misconception that Cloud Storage can serve as a queryable database for SQL, when in fact it requires an external query engine like BigQuery or Dataproc for SQL access.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Storage for raw data and BigQuery for processed data

Cloud Storage is the ideal service for storing raw, unstructured binary sensor data at petabyte scale, offering low-cost, durable object storage with multiple access tiers. BigQuery is a serverless, highly scalable data warehouse that allows data scientists to run SQL queries on processed data, with features like columnar storage and automatic optimization for analytical workloads. This combination directly addresses the need for raw storage and SQL-based analytics on processed data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Bigtable for raw data and Cloud Spanner for processed data

    Why it's wrong here

    Bigtable is a wide-column NoSQL store for low-latency key lookups, not binary object storage, and Spanner is a globally distributed relational database, not an analytics query engine over processed files. This pairing fits high-throughput transactional workloads, not a data lake.

  • ✗

    Cloud Storage for both raw and processed data

    Why it's wrong here

    Cloud Storage alone holds the raw binary files but provides no SQL engine, so data scientists cannot query processed data without an additional analytics service. It is the correct landing zone for raw objects, yet the stem also demands SQL querying capability.

  • ✗

    Cloud SQL for raw data and Cloud Dataproc for processing

    Why it's wrong here

    Cloud SQL is a managed relational database capped at modest storage, unsuitable for terabytes of raw binary sensor files, and Dataproc is a batch Spark/Hadoop service rather than an interactive SQL warehouse. This suits lift-and-shift relational workloads, not a data lake.

  • ✓

    Cloud Storage for raw data and BigQuery for processed data

    Why this is correct

    Cloud Storage holds the raw unstructured binary sensor files cost-effectively at terabyte scale, while BigQuery provides serverless SQL analytics over the processed data. This pairing separates cheap object storage from query compute, matching the differing access patterns described.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.