Databricks-DE-Pro Cost and Performance Optimization Practice Question
A data engineer is optimizing a Spark job that reads from a large Delta table and performs a join with a smaller dimension table. The job is running slowly, and the engineer suspects data skew and shuffle overhead are the main issues. Which two techniques should the engineer apply to improve performance and reduce cost? (Choose two.)
⚠ Common exam trap
The trap here is thinking that repartitioning or increasing shuffle partitions will solve skew, when in fact they can add overhead without addressing the root cause.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable adaptive query execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
Broadcast join eliminates shuffle for the smaller table, and enabling AQE with skew join optimization dynamically handles skew in the larger table. Together, they reduce shuffle overhead and mitigate stragglers, improving performance and lowering cost. Other options either introduce unnecessary shuffles or do not directly address skew.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of shuffle partitions to 2000.
Why it's wrong here
Increasing shuffle partitions can help with large data volumes but may create many small partitions, increasing overhead. It does not directly address skew; skewed keys will still cause uneven partition sizes. Without AQE or salting, this may not resolve the issue and could worsen performance due to task scheduling overhead. It is not a targeted solution for skew.
- ✗
Cache the larger table in memory before the join.
Why it's wrong here
Caching a large table may consume significant memory and cause spilling, potentially slowing down the job. It does not address the shuffle or skew issues. Caching is beneficial for iterative algorithms or repeated access, but for a single join, it may not provide a net benefit and could increase cost by requiring a larger cluster. It is not the most effective optimization here.
- ✗
Repartition the larger table on the join key before the join.
Why it's wrong here
Repartitioning the larger table on the join key can help with skew but also introduces a full shuffle, which is expensive. If the smaller table is broadcast, no shuffle is needed at all. Repartitioning is often used to mitigate skew, but it should be combined with other techniques like salting. In this scenario, broadcasting is more effective and avoids the shuffle entirely.
- ✓
Enable adaptive query execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
Why this is correct
Adaptive Query Execution (AQE) dynamically optimizes query plans at runtime, including handling skew by splitting skewed partitions. Enabling skew join optimization allows Spark to automatically detect and mitigate skew during joins, improving performance without manual intervention. This reduces shuffle overhead and prevents straggler tasks, leading to faster completion and lower cost.
- ✓
Use broadcast join for the smaller dimension table.
Why this is correct
Broadcast join sends the smaller table to all worker nodes, avoiding a shuffle of the larger table. This reduces network overhead and speeds up the join, especially when the smaller table fits in memory. It is a standard optimization for star-schema joins in Spark. By eliminating the shuffle, it also reduces disk I/O and CPU usage, leading to lower cost.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.