An engineer has discovered that a specific join operation is extremely slow due to severe data skew on the join key. Which strategy should they use to mitigate this?
Trap 1: Use the 'broadcast' hint on the skewed table.
Broadcasting a skewed table will likely lead to an OOM error because the skewed partition will exceed the memory capacity of the executor. If the skewed key is massive, the broadcast mechanism cannot handle the unbalanced distribution, making this a dangerous approach for high-skew scenarios.
Trap 2: Increase the number of worker nodes to improve parallelism.
Adding more nodes does not fix data skew because the bottleneck is the single task processing the skewed key. No matter how many nodes you add, that specific task remains on a single executor, and the job will still wait for that one worker to finish processing the skew.
- A
Use the 'broadcast' hint on the skewed table.
Why it fails: Broadcasting a skewed table will likely lead to an OOM error because the skewed partition will exceed the memory capacity of the executor. If the skewed key is massive, the broadcast mechanism cannot handle the unbalanced distribution, making this a dangerous approach for high-skew scenarios.
- B
Apply 'salting' to the join key.
Salting involves adding a random number to the join key in both tables, which breaks up the large, skewed partitions into smaller, more manageable pieces. This allows Spark to distribute the work evenly across multiple nodes, preventing the 'long-tail' task issue where one node processes the entire skew.
- C
Increase the number of worker nodes to improve parallelism.
Why it fails: Adding more nodes does not fix data skew because the bottleneck is the single task processing the skewed key. No matter how many nodes you add, that specific task remains on a single executor, and the job will still wait for that one worker to finish processing the skew.
- D
Enable AQE (Adaptive Query Execution).
While AQE can help with some forms of skew, it is not a guaranteed fix for all scenarios, especially severe data skew. Salting is a more explicit and robust method that ensures the data is physically distributed, whereas AQE might not be able to optimize the execution plan sufficiently.