Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?
⚠ Common exam trap
Candidates often suggest increasing cluster size to solve shuffling issues. While this helps performance, it does not reduce the *volume* of data shuffled; only filtering or broadcasting does that.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Broadcast join for small tables.
Reducing network traffic is critical in Spark performance tuning. By utilizing features like broadcast joins, you avoid the shuffle phase entirely when one table is small. Additionally, using filter and select operations early in the transformation pipeline ensures that only the necessary rows and columns are transmitted, significantly decreasing the total data volume shuffled during wide transformations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Broadcast join for small tables.
Why this is correct
Broadcast joins send the entire small table to every executor, allowing the join to occur locally. This eliminates the need to shuffle the large table, which is the most expensive part of a join, thereby reducing network overhead and significantly improving the performance of the overall job.
- ✗
Increase spark.sql.shuffle.partitions.
Why it's wrong here
Increasing shuffle partitions actually tends to increase the number of small files and metadata operations. While it might prevent OOM errors, it does not reduce the volume of data shuffled. Instead, it changes how that data is split, often leading to higher overhead due to more task scheduling.
- ✓
Filter and select columns early.
Why this is correct
Filtering rows and selecting specific columns early in the Spark pipeline minimizes the amount of data stored in memory and shuffled across the network. By reducing the width and length of the DataFrame before a shuffle occurs, you optimize the total volume of data processed by executors.
- ✗
Enable dynamic resource allocation.
Why it's wrong here
Dynamic allocation helps manage cluster resources by adding or removing executors based on workload. While it improves overall cluster efficiency and cost, it does not optimize the data flow within a single Spark job or reduce the actual volume of data transferred during shuffle operations.
- ✗
Increase spark.driver.memory.
Why it's wrong here
Driver memory is independent of the shuffle process performed by executors. While more driver memory is required for large result sets being collected, it has no impact on the volume of data being shuffled between executors during distributed operations like joins or group-by transformations.
Visual reference
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.