hardMultiple Select
PDE Practice Question: A company uses Cloud Dataproc for ephemeral…
A company uses Cloud Dataproc for ephemeral clusters to run batch jobs. They want to ensure job reliability and data quality. Which two configuration options should they use? (Choose two.)
⚠ Common exam trap
The trap here is that candidates might confuse cost-saving or performance features with reliability and data quality mechanisms. For example, enabling preemptible VMs (which are spot instances in Google Cloud) reduces cost but can cause job failures if workers are reclaimed. Idle timeout only deletes clusters after inactivity, not ensuring reliable job execution. Custom machine types improve performance but not reliability. In contrast, initialization actions for Google Cloud Dataproc ensure every ephemeral cluster node has the correct software and data sources, directly supporting job reliability and data quality. Graceful decommissioning allows workers to complete their tasks before being removed, preventing data loss during scaling down or cluster deletion. These two options directly address consistency and fault tolerance for Dataproc batch jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use initialization actions for cluster setup.
Option B is correct because initialization actions let you run scripts on every node at cluster creation, so you can install monitoring/validation agents, configure logging, and enforce consistent setup that supports job reliability and data quality across ephemeral clusters. Option E is correct because graceful decommissioning (set via yarn:yarn.resourcemanager.decommissioning.timeout or Dataproc's gracefulDecommissionTimeout) lets running tasks finish before workers are removed, preventing partial job failures and data corruption when nodes are scaled down or preempted. Option A is not appropriate here because preemptible VMs reduce cost but can be reclaimed at any time, which undermines reliability rather than improving it. Option C is not appropriate because idle timeout only deletes idle clusters to save cost; it does not improve job reliability or data quality. Option D is not appropriate because custom machine types tune CPU/memory performance but do not by themselves ensure reliable job completion or data quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable preemptible VMs for cost savings.
Why it's wrong here
Preemptible VMs are reclaimed within 24 hours, so executors vanish mid-job and tasks must rerun, undermining the reliability the scenario demands. They suit fault-tolerant, checkpointed or short-lived batch work where cost outweighs guaranteed completion, not jobs needing dependable output.
- ✓
Use initialization actions for cluster setup.
Why this is correct
Initialization actions run scripts on every node during cluster creation, installing agents, libraries or configuration needed for consistent job execution. This satisfies the reliability requirement by ensuring each ephemeral cluster starts identically, preventing job failures caused by missing dependencies or inconsistent setup.
- ✗
Enable idle timeout to automatically delete clusters.
Why it's wrong here
Idle timeout merely deletes a cluster after inactivity; it neither prevents task failure nor validates data, so reliability and quality remain unaddressed. It is intended to curb spend on forgotten ephemeral clusters, and would be the right pick for a cost-control question rather than this one.
- ✗
Use custom machine types for better performance.
Why it's wrong here
Custom machine types tune vCPU-to-memory ratios for a workload's resource profile, which affects performance and cost but not task retry behaviour or data validation. They belong in a sizing or tuning scenario, not one asking how to guarantee job reliability and data quality.
- ✓
Use graceful decommissioning of workers.
Why this is correct
Graceful decommissioning lets Dataproc drain running tasks from workers before removal, allowing in-flight work to finish and be rescheduled. This directly satisfies the reliability constraint for ephemeral clusters, preventing job failure or data loss when preemptible or scaled-down workers terminate mid-job.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.