Databricks-DE-Pro Developing Code (Python/SQL) Practice Question
You are debugging a PySpark job that is experiencing severe memory pressure during a join on a massive column. You suspect data skew. Which approach is best to mitigate this issue?
⚠ Common exam trap
Test-takers frequently suggest broadcasting the large skewed table or increasing executor memory, ignoring the root cause of uneven partition distribution.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement a salted join by adding a random integer to the join key.
Data skew is a common performance bottleneck in distributed systems where one task takes significantly longer than others because it processes a disproportionate amount of data. By salting the skewed key—adding a random prefix or suffix—the data is distributed more evenly across the cluster partitions. This technique prevents executor memory overflow and allows for parallel processing of the skewed key, significantly improving overall job performance and reliability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of partitions using 'spark.sql.shuffle.partitions'.
Why it's wrong here
Increasing partitions helps with small-file problems but does not solve skew. If all data for a specific skewed key is mapped to the same partition, adding more partitions will not redistribute the load. The skewed key will still end up on a single task, causing the same memory pressure.
- ✓
Implement a salted join by adding a random integer to the join key.
Why this is correct
Salting splits a single skewed key into multiple keys by appending a random integer. This spreads the data across multiple partitions, ensuring that no single task is overwhelmed. By expanding the join key on both sides, the join can be completed in parallel, effectively resolving the performance bottleneck.
- ✗
Use the 'cache()' method on the skewed dataframe.
Why it's wrong here
Caching the dataframe persists it in the executor memory, but it does not address the skew itself. If the data is skewed, the partition holding the skewed key will still be too large, leading to out-of-memory errors regardless of whether the dataframe is cached or processed on-the-fly.
- ✗
Enable AQE and increase the 'spark.sql.adaptive.skewJoin.enabled' setting.
Why it's wrong here
While AQE can optimize some skew, it is not a universal solution for extreme data skew. Relying solely on configuration settings without understanding the distribution of data is risky. Salting provides a deterministic, developer-controlled mitigation strategy that is more robust than relying on the optimizer's heuristics alone.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.