Courseiva

PDE Designing Data Processing Systems Practice Question

A financial services firm needs to design a batch processing system on Google Cloud to analyze large volumes of historical transaction data stored in Cloud Storage. The data is in Parquet format and must be processed using Apache Spark. The firm wants to minimize operational overhead and only pay for the resources used during job execution. Which Google Cloud service should they use?

⚠ Common exam trap

Candidates often confuse Dataflow with a Spark execution service or assuming that Dataproc requires a persistent cluster, when serverless Spark is available.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataproc Serverless for Spark

Cloud Dataproc Serverless for Spark provides a fully managed Spark environment without cluster provisioning. It charges only for the resources used during job execution, aligning with the cost and operational requirements. It supports reading Parquet from Cloud Storage and is designed for batch processing workloads, making it the best fit for analyzing historical transaction data with Spark while minimizing overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Dataflow with a custom Apache Spark runner

    Why it's wrong here

    Cloud Dataflow is a fully managed service for Apache Beam, not a native Spark execution environment. While it can run pipelines written with the Beam API, it does not natively execute Apache Spark jobs. Using a custom Spark runner on Dataflow is not a supported or standard approach, and it would introduce complexity and potential compatibility issues, failing to meet the requirement for a straightforward Spark processing solution.

  • ✓

    Cloud Dataproc Serverless for Spark

    Why this is correct

    Cloud Dataproc Serverless for Spark is a fully managed, serverless Spark environment that eliminates the need to provision or manage clusters. It automatically scales resources based on workload and charges only for the resources consumed during job execution. It natively supports Spark jobs and can read Parquet data from Cloud Storage, making it the ideal choice for minimizing operational overhead and paying only for what is used.

  • ✗

    BigQuery with external tables over Cloud Storage

    Why it's wrong here

    BigQuery is a serverless analytics warehouse that can query external data in Cloud Storage, but it does not execute Apache Spark jobs. While it can process Parquet data, it uses its own SQL engine, not Spark. This does not meet the requirement to use Apache Spark for processing, and it may not be suitable for complex Spark transformations. Therefore, it is not the correct service for this scenario.

  • ✗

    Cloud Dataproc with a long-running cluster

    Why it's wrong here

    A long-running Dataproc cluster incurs costs even when no jobs are running, which violates the requirement to only pay for resources used during job execution. While Dataproc supports Spark and Parquet, the continuous operation of the cluster adds operational overhead and unnecessary expense. This approach does not minimize operational overhead or align with the cost model of paying only during job execution.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.