PDE Designing Data Processing Systems Practice Question
A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?
⚠ Common exam trap
The trap here is assuming that data must be exported from BigQuery to Cloud Storage for Spark processing, overlooking the direct BigQuery connector available in Dataproc Serverless.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.
Dataproc Serverless for Spark with the BigQuery connector allows the company to run Spark jobs directly on BigQuery data without exporting it. It is serverless, so it minimizes operational overhead and costs by charging only for job execution. It supports complex Spark MLlib models and iterative processing, making it the best fit for weekly training on historical GPS data. This approach aligns with the goals of minimizing cost and operational overhead while leveraging Spark's capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Run the Spark job on a long-running Dataproc cluster that is always available.
Why it's wrong here
A long-running cluster incurs costs even when idle, which violates the requirement to minimize cost. It also requires ongoing management and patching. While it can run Spark jobs and read from BigQuery, the continuous operation is not cost-effective for a weekly training job. This approach does not minimize operational overhead or cost, as the cluster runs 24/7 regardless of job frequency.
- ✓
Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.
Why this is correct
Dataproc Serverless for Spark can read data directly from BigQuery using the BigQuery connector, eliminating the need to export data. It is a serverless solution, so it automatically provisions and scales resources, and charges only for the duration of the job. This minimizes both cost and operational overhead, as there is no cluster to manage. It supports complex Spark MLlib models and iterative processing, making it ideal for this scenario.
- ✗
Use BigQuery ML to train the model directly in BigQuery without Spark.
Why it's wrong here
BigQuery ML supports certain model types but may not support the complex Spark MLlib model required. It also may not handle iterative processing over the entire dataset in the same way as Spark. If the model requires custom Spark code, BigQuery ML is not suitable. This approach does not meet the requirement to use Apache Spark for the complex machine learning model, and it may lack the flexibility needed for iterative training.
- ✗
Export the data to Cloud Storage and run a Dataproc cluster with autoscaling, then delete the cluster after training.
Why it's wrong here
Exporting data to Cloud Storage adds an extra step and may incur additional time and cost. While a Dataproc cluster with autoscaling can handle Spark jobs, it still requires cluster management and provisioning. Deleting the cluster after training reduces cost but does not eliminate the operational overhead of cluster creation and configuration. This approach is less efficient than using a serverless option that directly integrates with BigQuery.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.