DEA-C02 Performance Optimization Practice Question
A data engineer is optimizing a query that joins a very large fact table to a small dimension table. The Query Profile shows the small table being redistributed across all nodes before the join. Which action is most likely to improve performance?
⚠ Common exam trap
The trap here is assuming that any redistribution in a join is unavoidable and must simply be run faster with a larger warehouse, when the distribution strategy itself can be changed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Broadcast the small dimension table so it is replicated to all nodes and the large fact table is not redistributed.
The profile shows the small relation being shuffled, which is the expensive but avoidable side of the join. Broadcasting the small table replicates it to every node so the large fact table stays in place, eliminating the shuffle of the big input. This is the standard remedy when a join's redistribution targets the wrong relation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Broadcast the small dimension table so it is replicated to all nodes and the large fact table is not redistributed.
Why this is correct
When one side of a join is small, broadcasting it to every node lets each node join its local fact rows without shuffling the large table. This converts a costly redistribution of the large input into a cheap replication of the small input, which is exactly the pattern the profile is hinting at as the bottleneck.
- ✗
Increase the warehouse size so the redistribution of the small table completes faster across more nodes.
Why it's wrong here
Adding nodes increases the number of destinations the small table must be sent to, which can make the broadcast-style redistribution more expensive rather than cheaper. The choice of join distribution strategy, not raw compute, governs this cost, so scaling the warehouse does not correct the plan the optimizer selected.
- ✗
Rewrite the query to use a correlated subquery instead of a join so the small table is not redistributed.
Why it's wrong here
Correlated subqueries are typically executed as joins or per-row lookups internally and often perform worse than a clean join. Replacing a set-based join with row-by-row semantics does not eliminate the redistribution the optimizer chose and usually increases execution time, making this a step backward rather than an optimization.
- ✗
Add a clustering key to the small dimension table so its rows are stored contiguously.
Why it's wrong here
Clustering improves pruning for filtered scans, but the dimension table is already small and fully read, so ordering its rows does not reduce the redistribution the profile shows. The expensive step is moving the small table's rows to every node, which clustering cannot prevent, so this action adds maintenance cost without touching the bottleneck.
About these practice questions
This DEA-C02 question is part of Courseiva's 229-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Snowflake exam blueprint
This DEA-C02 practice question is part of Courseiva's free Snowflake certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DEA-C02 exam.