Courseiva

PDE Designing Data Processing Systems Practice Question

A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?

⚠ Common exam trap

The trap here is assuming that data must be exported from BigQuery to Cloud Storage for Spark processing, overlooking the direct BigQuery connector available in Dataproc Serverless.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.

Dataproc Serverless for Spark with the BigQuery connector allows the company to run Spark jobs directly on BigQuery data without exporting it. It is serverless, so it minimizes operational overhead and costs by charging only for job execution. It supports complex Spark MLlib models and iterative processing, making it the best fit for weekly training on historical GPS data. This approach aligns with the goals of minimizing cost and operational overhead while leveraging Spark's capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Run the Spark job on a long-running Dataproc cluster that is always available.

    Why it's wrong here

    A long-running cluster incurs costs even when idle, which violates the requirement to minimize cost. It also requires ongoing management and patching. While it can run Spark jobs and read from BigQuery, the continuous operation is not cost-effective for a weekly training job. This approach does not minimize operational overhead or cost, as the cluster runs 24/7 regardless of job frequency.

  • ✓

    Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.

    Why this is correct

    Dataproc Serverless for Spark can read data directly from BigQuery using the BigQuery connector, eliminating the need to export data. It is a serverless solution, so it automatically provisions and scales resources, and charges only for the duration of the job. This minimizes both cost and operational overhead, as there is no cluster to manage. It supports complex Spark MLlib models and iterative processing, making it ideal for this scenario.

  • ✗

    Use BigQuery ML to train the model directly in BigQuery without Spark.

    Why it's wrong here

    BigQuery ML supports certain model types but may not support the complex Spark MLlib model required. It also may not handle iterative processing over the entire dataset in the same way as Spark. If the model requires custom Spark code, BigQuery ML is not suitable. This approach does not meet the requirement to use Apache Spark for the complex machine learning model, and it may lack the flexibility needed for iterative training.

  • ✗

    Export the data to Cloud Storage and run a Dataproc cluster with autoscaling, then delete the cluster after training.

    Why it's wrong here

    Exporting data to Cloud Storage adds an extra step and may incur additional time and cost. While a Dataproc cluster with autoscaling can handle Spark jobs, it still requires cluster management and provisioning. Deleting the cluster after training reduces cost but does not eliminate the operational overhead of cluster creation and configuration. This approach is less efficient than using a serverless option that directly integrates with BigQuery.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.