PDE Designing Data Processing Systems Practice Question
A financial services firm needs to design a batch processing system on Google Cloud to analyze large volumes of historical transaction data stored in Cloud Storage. The data is in Parquet format and must be processed using Apache Spark. The firm wants to minimize operational overhead and only pay for the resources used during job execution. Which Google Cloud service should they use?
⚠ Common exam trap
Candidates often confuse Dataflow with a Spark execution service or assuming that Dataproc requires a persistent cluster, when serverless Spark is available.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataproc Serverless for Spark
Cloud Dataproc Serverless for Spark provides a fully managed Spark environment without cluster provisioning. It charges only for the resources used during job execution, aligning with the cost and operational requirements. It supports reading Parquet from Cloud Storage and is designed for batch processing workloads, making it the best fit for analyzing historical transaction data with Spark while minimizing overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Dataflow with a custom Apache Spark runner
Why it's wrong here
Cloud Dataflow is a fully managed service for Apache Beam, not a native Spark execution environment. While it can run pipelines written with the Beam API, it does not natively execute Apache Spark jobs. Using a custom Spark runner on Dataflow is not a supported or standard approach, and it would introduce complexity and potential compatibility issues, failing to meet the requirement for a straightforward Spark processing solution.
- ✓
Cloud Dataproc Serverless for Spark
Why this is correct
Cloud Dataproc Serverless for Spark is a fully managed, serverless Spark environment that eliminates the need to provision or manage clusters. It automatically scales resources based on workload and charges only for the resources consumed during job execution. It natively supports Spark jobs and can read Parquet data from Cloud Storage, making it the ideal choice for minimizing operational overhead and paying only for what is used.
- ✗
BigQuery with external tables over Cloud Storage
Why it's wrong here
BigQuery is a serverless analytics warehouse that can query external data in Cloud Storage, but it does not execute Apache Spark jobs. While it can process Parquet data, it uses its own SQL engine, not Spark. This does not meet the requirement to use Apache Spark for processing, and it may not be suitable for complex Spark transformations. Therefore, it is not the correct service for this scenario.
- ✗
Cloud Dataproc with a long-running cluster
Why it's wrong here
A long-running Dataproc cluster incurs costs even when no jobs are running, which violates the requirement to only pay for resources used during job execution. While Dataproc supports Spark and Parquet, the continuous operation of the cluster adds operational overhead and unnecessary expense. This approach does not minimize operational overhead or align with the cost model of paying only during job execution.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.