easyMultiple Choice
PDE Practice Question: A data engineer needs to process large CSV files…
A data engineer needs to process large CSV files (hundreds of GB) stored in Cloud Storage using Spark on a Dataproc cluster. The job performs a series of transformations and aggregations. Which configuration is most cost-effective and operationally efficient?
⚠ Common exam trap
Google Cloud often tests the misconception that preemptible VMs are unreliable for all workloads, but in Spark batch processing with fault tolerance, they are both cost-effective and operationally efficient, unlike stateful or latency-sensitive applications.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a cluster with a standard master node and 10 preemptible worker nodes (n1-standard-4).
Preemptible workers are significantly cheaper (about 80% discount) and ideal for batch processing of large CSV files where fault tolerance is built into Spark via RDD lineage. Using standard nodes for the master ensures cluster stability, while preemptible workers handle the distributed transformations and aggregations cost-effectively. This configuration balances cost and operational efficiency for ephemeral, fault-tolerant workloads.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a cluster with 10 high-memory (n1-highmem-8) VMs as workers to improve shuffle performance.
Why it's wrong here
Highmem machine types cost more per vCPU, and shuffle-heavy Spark aggregations are constrained by CPU and disk rather than memory alone, so the premium is wasted. Highmem workers fit memory-bound workloads such as large in-memory caches or wide joins that spill under standard memory.
- ✓
Use a cluster with a standard master node and 10 preemptible worker nodes (n1-standard-4).
Why this is correct
Preemptible workers cost far less than standard VMs, and ten n1-standard-4 nodes provide enough parallelism for hundreds of GB of CSV transformations. A standard master preserves cluster stability, making this the cost-effective, operationally efficient configuration.
- ✗
Use a single-node cluster with a high-memory machine type.
Why it's wrong here
A single node cannot parallelise Spark transformations across hundreds of GB, so all shuffle and aggregation work serialises on one machine, extending runtime and cost. Single-node clusters suit small datasets or development testing, where cluster startup overhead outweighs distribution benefits.
- ✗
Use a cluster with 10 standard (n1-standard-4) VMs as master and worker nodes, all non-preemptible.
Why it's wrong here
Ten non-preemptible n1-standard-4 nodes pay full price for capacity that preemptible workers would supply far cheaper, and standard memory may throttle shuffle-heavy aggregations. Non-preemptible uniform clusters fit latency-sensitive or checkpoint-free jobs where worker loss is unacceptable.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.