Courseiva
Debugging and Deploying →mediumMultiple Choice

Databricks-DE-Pro Debugging and Deploying Practice Question

A data engineer is debugging a slow-running query. They notice that the data is skewed, causing one task to take significantly longer than others. Which approach effectively addresses this skew?

⚠ Common exam trap

Candidates often suggest simply increasing the cluster node count or core count, failing to realize that data skew leaves specific executors idle while one overloaded task bottlenecks the entire stage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Add a salt column to the skewed key to distribute the data across more partitions.

Data skew occurs when one partition contains disproportionately more data than others, causing a single executor to bottleneck the entire stage. By using techniques like salt, broadcast joins, or repartitioning, the engineer can distribute the load more evenly across the cluster. This is essential for optimizing performance and preventing timeouts, ensuring that production jobs adhere to defined SLAs and utilize cluster resources efficiently.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the cluster's disk size to allow for more local shuffle storage.

    Why it's wrong here

    Increasing disk size does not solve the underlying problem of uneven data distribution. The bottleneck remains at the CPU and memory level of the overloaded executor. The job will still suffer from the same performance degradation because the data processing load is not being shared among other executors.

  • ✓

    Add a salt column to the skewed key to distribute the data across more partitions.

    Why this is correct

    Adding a salt to the join or grouping key forces Spark to distribute the data evenly across partitions. By breaking up the massive partition into smaller, manageable chunks, the workload becomes balanced, significantly reducing the execution time of the stage that was previously suffering from the data skew bottleneck.

  • ✗

    Change the file format from Parquet to CSV to reduce overhead.

    Why it's wrong here

    CSV is a text-based, uncompressed format that is significantly less efficient than Parquet. Switching to CSV would increase I/O overhead and memory consumption, potentially making the skew issue worse. Parquet's columnar structure is optimized for performance, and it is not the cause of data skew in join operations.

  • ✗

    Enable 'Auto-scaling' to automatically add more nodes during the skewed stage.

    Why it's wrong here

    Auto-scaling adds nodes to the cluster, but it cannot re-partition data that is already being processed by a single task. Even with more nodes, the skewed partition will still be assigned to one executor, leaving other nodes idle while waiting for that single, overloaded task to finish processing.

About these practice questions

Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.