Courseiva

PDE Designing Data Processing Systems Practice Question

A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?

⚠ Common exam trap

The trap is assuming 'reduce cost' means 'reduce cluster size' — candidates pick single-node or fewer workers, but the exam expects you to recognize that preemptible/Spot instances reduce cost per unit while preserving job characteristics.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use preemptible instances for worker nodes

Preemptible (Spot) VMs cost up to 80% less than standard VMs and are ideal for fault-tolerant, batch-oriented workloads like Spark ML jobs that can tolerate occasional preemption. Since the jobs run only 2 hours daily and Dataproc automatically handles node replacement when preemptible instances are reclaimed, using preemptible workers delivers the largest cost reduction without changing job characteristics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a single-node cluster to eliminate overhead

    Why it's wrong here

    A single-node cluster runs master and worker roles on one VM, capping the cores and memory available to Spark ML and preventing the distributed parallelism the jobs rely on. It is tempting because it removes multi-VM overhead, and it would be correct for small, non-distributed workloads.

  • ✗

    Enable high-availability mode to avoid restarts

    Why it's wrong here

    High-availability mode runs extra master nodes and a second worker replica, raising the hourly rate for a two-hour batch job that does not need fault tolerance. It is tempting because it prevents restart delays, and it would be correct for long-running streaming clusters where a master failure would lose in-flight work.

  • ✓

    Use preemptible instances for worker nodes

    Why this is correct

    Preemptible instances cost significantly less than standard Dataproc worker nodes, and Spark can tolerate their loss through retries and task rescheduling. Since the daily jobs are short and their characteristics stay unchanged, using preemptible workers cuts compute spend without altering the job design.

  • ✗

    Increase the number of standard workers to finish faster

    Why it's wrong here

    Adding standard workers raises the per-second cost of the cluster while the two-hour runtime is set by the job's data volume and algorithm, so spend increases without a proportional time saving. It is tempting because faster completion reduces billed minutes, and it would be correct for a deadline-bound job that scales linearly.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.