easyMultiple Select
PDE Practice Question: Which TWO actions can reduce the cost of running…
Which TWO actions can reduce the cost of running a Dataproc cluster for a nightly batch job?
⚠ Common exam trap
Google Cloud often tests the misconception that scaling up resources (more nodes or faster hardware) always reduces cost by shortening runtime, but in reality, the increased per-hour cost usually outweighs the time savings for batch jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use preemptible VMs for worker nodes.
Option C is correct because preemptible VMs in Dataproc cost significantly less than standard VMs (up to 80% cheaper), and for a nightly batch job that can tolerate interruptions, using preemptible worker nodes reduces compute costs substantially. Option E is correct because Dataproc charges for cluster resources while the cluster exists; deleting the cluster after the nightly job completes ensures you only pay for the duration of the job rather than leaving idle nodes running. Option A is incorrect because increasing worker nodes raises the number of billable VM instances, increasing cost rather than reducing it. Option B is incorrect because high-memory machine types for the master node increase the per-hour price of that node without benefiting a typical batch workload. Option D is incorrect because attaching local SSDs adds cost (and local SSD pricing) to every node, increasing rather than reducing the cluster's expense.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of worker nodes for faster processing.
Why it's wrong here
Adding worker nodes increases the total vCPU and memory hours billed, so the cluster costs more even if the job finishes sooner. Scaling workers suits deadline-bound jobs where throughput matters, not a nightly batch where cost is the constraint.
- ✗
Use high-memory machine types for master node.
Why it's wrong here
High-memory master nodes raise the per-hour rate without speeding the nightly batch, since the master only coordinates YARN and HDFS metadata rather than processing data. Such machine types suit memory-intensive driver workloads, not cost reduction on a fixed nightly schedule.
- ✓
Use preemptible VMs for worker nodes.
Why this is correct
Preemptible VMs cost substantially less than standard worker instances, directly cutting the dominant compute expense of the nightly batch job. Because Dataproc tolerates worker preemption by re-running affected tasks, the job's fault-tolerant batch nature satisfies the availability constraint that would otherwise rule this option out.
- ✗
Attach local SSDs to all nodes.
Why it's wrong here
Local SSDs are ephemeral and billed separately from persistent disk, adding cost while providing no benefit for a batch job whose data already resides in Cloud Storage. They suit shuffle-heavy or latency-sensitive workloads requiring fast scratch storage.
- ✓
Delete the cluster after the job completes.
Why this is correct
Deleting the cluster after the job completes eliminates charges for idle compute and storage once the nightly batch finishes, since Dataproc bills per-minute for the cluster's lifetime. This directly satisfies the cost-reduction constraint by ensuring no resources persist between runs, unlike a long-lived cluster that accrues charges continuously.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.