Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

Your team uses Cloud Dataproc for Spark ML training jobs. You want to reduce costs for non-critical, fault-tolerant training jobs. Which Dataproc feature should you use for worker nodes?

⚠ Common exam trap

A common pitfall is confusing committed use discounts (which require a 1- or 3-year commitment) with preemptible VMs (which are interruptible but cost-effective). For non-critical, fault-tolerant workloads, preemptible VMs are the appropriate cost-saving feature, not committed use discounts.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use preemptible instances for worker nodes.

Preemptible instances are short-lived, lower-cost VMs that Cloud Dataproc can use for worker nodes. Because the training jobs are non-critical and fault-tolerant (e.g., they can handle node failures via Spark's built-in resilience), preemptible instances significantly reduce costs while still completing the workload. This directly addresses the requirement to reduce costs for fault-tolerant jobs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use preemptible instances for worker nodes.

    Why this is correct

    Preemptible instances suit fault-tolerant Spark training because Dataproc replaces them automatically when Compute Engine reclaims capacity, and they cost substantially less than standard VMs. The stem's non-critical, fault-tolerant constraint is satisfied: interrupted workers are re-added without failing the job, unlike sole-primary or non-retryable workloads.

  • ✗

    Use custom machine types with more memory.

    Why it's wrong here

    Larger custom machine types raise per-hour compute cost, the opposite of the goal. They suit memory-intensive workloads needing more RAM per core, not fault-tolerant batch jobs where workers can be reclaimed. Cost reduction here comes from preemptible or Spot workers, which are cheaper but interruptible.

  • ✗

    Use SSDs instead of HDDs for persistent disks.

    Why it's wrong here

    SSD persistent disks cost more per GB than standard HDD, increasing spend rather than reducing it. SSDs are the right pick when workloads are I/O-bound and need higher throughput or lower latency. For non-critical, fault-tolerant training, cheaper interruptible workers address cost, not disk type.

  • ✗

    Use committed use discounts for 1-year or 3-year terms.

    Why it's wrong here

    Committed use discounts cut cost only by locking in one- or three-year spend, which suits steady, predictable, long-running workloads. Non-critical fault-tolerant training is intermittent and reclaimable, so preemptible or Spot workers deliver the saving without a long commitment. CUDs also cannot apply to preemptible instances.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.