Databricks-DA-Assoc Analyzing Queries Practice Question
A data analyst runs a query in Databricks SQL that returns the total sales per region. The query is slow, and the analyst notices that the execution plan shows a full table scan on a large Delta table. The analyst wants to reduce the amount of data read. Which action is most likely to improve performance?
⚠ Common exam trap
The trap here is focusing on cluster resources or caching instead of addressing the fundamental issue: reading unnecessary data due to lack of partition pruning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a WHERE clause on a partitioned column to enable partition pruning.
Partition pruning is the most direct way to reduce data read. By filtering on a partitioned column, Delta Lake can skip partitions that do not match the filter, turning a full table scan into a targeted read. Other options do not reduce the volume of data read or are inapplicable to an aggregation query.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a broadcast hint to force a broadcast join.
Why it's wrong here
The query is an aggregation, not a join, so a broadcast hint is irrelevant. Even if it were a join, broadcasting would not reduce the scan of the large table; it would only affect the join strategy. The root cause is reading all data, not the join method.
- ✗
Cache the table using CACHE TABLE before running the query.
Why it's wrong here
Caching stores data in memory, which can speed up repeated queries, but it does not reduce the initial read of all data. For a one-time query, caching adds overhead and does not address the full table scan. It is not a substitute for partition pruning.
- ✓
Add a WHERE clause on a partitioned column to enable partition pruning.
Why this is correct
If the table is partitioned on a column used in a filter, adding a WHERE clause on that column allows Delta to skip entire partitions. This reduces the data read from the full table scan to only relevant partitions, directly addressing the slow performance caused by scanning all data.
- ✗
Increase the cluster size to add more worker nodes.
Why it's wrong here
Adding more workers increases parallel processing power, but if the query is reading all data due to a full table scan, the I/O bottleneck remains. More workers may not help if the data is not pruned; the query still reads the same volume, and performance gains may be marginal.
About these practice questions
One of 291 original Databricks-DA-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.