easyMultiple Choice
PDE Practice Question: A large retail company processes point-of-sale…
A large retail company processes point-of-sale transactions from thousands of stores daily. The current batch pipeline runs on Cloud Dataproc using Spark and takes 3 hours to complete. The business wants to reduce processing time to under 30 minutes. The pipeline reads from Cloud Storage, joins with inventory data from BigQuery, performs aggregations, and writes to Cloud SQL for reporting. What is the most effective optimization?
⚠ Common exam trap
Candidates often assume that simply scaling up the existing infrastructure (more workers or auto-scaling) is the most effective optimization, but Cisco tests the understanding that architectural changes to reduce data movement and leverage service-specific strengths (like BigQuery for joins) are far more impactful than brute-force scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Read inventory data from BigQuery and pre-join in BigQuery, then export to Cloud Storage as ORC files
It offloads the join operation to BigQuery, which is optimized for large-scale analytics and can process the join much faster than Spark. By pre-joining and exporting the result as ORC files (a columnar format optimized for Spark), the pipeline avoids the expensive shuffle and data transfer between Cloud Storage and BigQuery, significantly reducing the overall processing time to meet the 30-minute target.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Migrate the pipeline to Cloud Dataflow with Apache Beam for auto-scaling
Why it's wrong here
Effective but requires significant code changes; might be overkill.
- ✓
Read inventory data from BigQuery and pre-join in BigQuery, then export to Cloud Storage as ORC files
Why this is correct
Reduces data shuffle in Spark and speeds up processing.
- ✗
Write intermediate results to Cloud SQL instead of BigQuery for faster access
Why it's wrong here
Cloud SQL is not designed for large-scale analytical joins; would bottleneck.
- ✗
Increase the number of worker nodes in the Dataproc cluster
Why it's wrong here
Scalability helps but may not achieve 30 minutes; also higher cost.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.