Courseiva

Databricks-DE-Pro Developing Code (Python/SQL) Practice Question

You are optimizing a PySpark job that reads from a Delta table. You notice skewed data distribution on the 'customer_id' column, causing Task-level stragglers. Which transformation should you apply to the DataFrame to mitigate this skew during a join operation?

⚠ Common exam trap

Candidates often confuse salting with partitioning. They mistakenly attempt to repartition the DataFrame using the skewed key itself, which does not address the uneven distribution of records across partitions during the shuffle phase.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Salt the skewed column and perform the join.

Salting involves adding a random prefix to the join key of the skewed table and exploding the join key of the dimension table. This technique breaks down massive partitions caused by skewed keys into smaller, evenly distributed chunks. Understanding data skew is critical in Databricks for maintaining cluster stability and preventing OOM errors, ensuring that shuffle partitions are utilized efficiently across all available executors rather than overloading one core.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the spark.sql.shuffle.partitions configuration dynamically.

    Why it's wrong here

    Increasing partition count spreads data further but does not solve the underlying issue where one partition contains the vast majority of data. The executor processing that specific partition will still take significantly longer than others, failing to address the fundamental imbalance caused by highly frequent keys.

  • ✗

    Broadcast the skewed table to all executors.

    Why it's wrong here

    Broadcasting requires the table to fit entirely within the memory of the driver and executors. If the skewed table is large enough to cause partitioning issues, it will likely exceed memory limits when broadcast, leading to immediate job failure through OutOfMemory errors during the broadcast process.

  • ✓

    Salt the skewed column and perform the join.

    Why this is correct

    Salting distributes the rows associated with the skewed key across multiple partitions by appending a random integer. This forces the join operation to process these rows in parallel across different executors, effectively eliminating the bottleneck caused by the skewed distribution of the customer_id column during the shuffle phase.

  • ✗

    Cache the skewed table in memory before the join.

    Why it's wrong here

    Caching keeps data in memory but does not reorganize the distribution of that data across tasks. Even if the data is cached, the join algorithm still relies on the original partitioning scheme, meaning the skew remains present and will continue to impact task execution time performance significantly.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.