PDE Designing Data Processing Systems Practice Question
A startup wants to analyze user clickstream data stored in Cloud Storage in Parquet format. They need to run ad-hoc SQL queries without managing any servers and want to pay only for the queries they run. Which Google Cloud service should they use?
⚠ Common exam trap
The trap here is assuming that any data processing service can query Parquet files, but only BigQuery offers serverless SQL with external table support.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BigQuery
BigQuery provides serverless SQL analytics with pay-per-query pricing and can query external Parquet data in Cloud Storage. It requires no infrastructure management, perfectly matching the startup's needs. Other services either require provisioning, are not designed for SQL analytics, or lack external data querying.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
BigQuery
Why this is correct
BigQuery is a serverless, highly scalable data warehouse that supports querying external data in Cloud Storage via external tables. It charges based on the amount of data scanned per query, aligning with pay-per-query. It requires no infrastructure management, making it ideal for ad-hoc SQL analysis.
- ✗
Dataproc
Why it's wrong here
Dataproc is a managed Spark and Hadoop service that requires cluster provisioning and management, even if ephemeral. It is not serverless and does not offer pay-per-query pricing; costs are based on cluster usage. It is overkill for simple ad-hoc SQL queries.
- ✗
Cloud Bigtable
Why it's wrong here
Bigtable is a NoSQL database for high-throughput, low-latency workloads, not for SQL analytics on Parquet files. It requires schema design and does not support ad-hoc SQL queries directly on external data, making it inappropriate for this use case.
- ✗
Cloud SQL
Why it's wrong here
Cloud SQL is a managed relational database service for OLTP workloads, not for analyzing large Parquet files in Cloud Storage. It requires provisioning instances and does not natively query external files, making it unsuitable for ad-hoc analysis of clickstream data.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.