easyMultiple Choice
PDE Practice Question: A company runs a batch ETL pipeline on Cloud…
A company runs a batch ETL pipeline on Cloud Dataproc. During peak hours, the job takes longer than expected. The pipeline reads from Cloud Storage, transforms data, and writes to BigQuery. What is the most cost-effective way to improve performance without redesigning the pipeline?
⚠ Common exam trap
Candidates often assume scaling up the master node or improving local storage will help, but the exam tests understanding that horizontal scaling with cheap, ephemeral workers is the most cost-effective approach for batch processing workloads that are CPU-bound and fault-tolerant.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a secondary worker group using preemptible VMs and increase the number of workers.
Adding a secondary worker group with preemptible VMs is the most cost-effective way to improve performance because it allows you to scale out the cluster horizontally with compute instances that are significantly cheaper (up to 80% discount) than regular VMs. This directly addresses the bottleneck of processing capacity during peak hours without requiring any pipeline redesign, as Cloud Dataproc can automatically distribute work across additional workers.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Add a secondary worker group using preemptible VMs and increase the number of workers.
Why this is correct
Preemptible secondary workers cost far less than standard instances, and adding workers increases parallel processing capacity for the Cloud Storage to BigQuery job. This improves throughput during peak hours without altering the pipeline design, satisfying the cost-effectiveness and no-redesign constraints.
- ✗
Enable local SSDs on all worker nodes.
Why it's wrong here
SSDs increase cost, and the bottleneck is likely CPU or memory.
- ✗
Increase the master node's machine type to n1-highmem-32.
Why it's wrong here
Master node size does not affect worker parallelism.
- ✗
Use Cloud Composer to schedule the job with a higher priority.
Why it's wrong here
Composer does not improve runtime of the job itself.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.