Databricks-Spark-Assoc Spark Architecture and Components Practice Question
A Databricks engineer is diagnosing why a Spark job's shuffle phase writes a very large amount of data to disk. The engineer wants to reduce shuffle overhead by changing how the job is structured and configured. Which TWO actions are most likely to reduce the volume of shuffle data written? (Choose two.)
⚠ Common exam trap
The trap here is treating a larger shuffle partition count as a way to reduce shuffle data, when it only splits the same volume into more files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Pre-aggregate data with reduceByKey before a subsequent join so fewer records participate in the shuffle.
Eliminating a shuffle through broadcast hash join and reducing record counts through map-side pre-aggregation both cut the actual bytes that must be written and read across the network. Partition-count tuning, serializer choice, and caching change performance characteristics without removing the underlying data movement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase spark.sql.shuffle.partitions to a much larger value without changing the query plan.
Why it's wrong here
Raising the shuffle partition count creates more, smaller output files per task but does not reduce the total bytes shuffled. The same records still cross the network, so shuffle write volume stays essentially the same while adding scheduling overhead and more small files.
- ✓
Pre-aggregate data with reduceByKey before a subsequent join so fewer records participate in the shuffle.
Why this is correct
Combining values locally with reduceByKey before the shuffle reduces the number of records that must be repartitioned, lowering both shuffle write and read volume. This map-side aggregation is the classic optimization for reducing network traffic in wide transformations.
- ✓
Use broadcast hash join instead of sort-merge join when one side of the join is small enough to fit within the broadcast threshold.
Why this is correct
A broadcast hash join ships the small table to every executor and avoids repartitioning the large table across the network, eliminating the shuffle entirely for that join. This directly reduces shuffle write volume, which is the metric the engineer is investigating in the Spark UI's shuffle stage details.
- ✗
Call persist(MEMORY_ONLY) on the DataFrame before the shuffle stage.
Why it's wrong here
Caching a DataFrame in memory speeds up repeated access but does not alter the number of records repartitioned during a shuffle. The shuffle still writes the same data to disk, so caching does not address the root cause of the large shuffle write volume described.
- ✗
Enable the Kryo serializer instead of the default Java serializer for the shuffle.
Why it's wrong here
Kryo reduces the serialized size of each record, which can shrink shuffle bytes somewhat, but it does not change how many records are shuffled. The question asks for actions most likely to reduce shuffle volume structurally, and serialization tuning is a secondary, smaller-effect optimization compared to eliminating or pre-aggregating the shuffle.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.