Courseiva

Databricks-DA-Assoc Executing Queries with Databricks SQL Practice Question

Exhibit

SELECT * FROM sales WHERE order_date >= '2023-01-01' AND region = 'North';

Refer to the exhibit. The table 'sales' is partitioned by 'region' and 'order_date'. How does the Databricks SQL engine process this query to ensure optimal performance?

⚠ Common exam trap

Candidates mistakenly think the engine scans all directories and filters rows afterward, ignoring how directory-level partition pruning works in Delta Lake.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It uses partition pruning to skip files in partitions not matching the filters

Databricks SQL utilizes partition pruning to ignore directories that do not match the WHERE clause filters. By identifying that the query filters on both 'region' and 'order_date', the engine restricts the scan to specific subdirectories on storage. This significantly reduces the amount of I/O required, which is the primary driver of performance in large-scale data lake queries, as it avoids reading irrelevant data from the storage layer.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It performs a full table scan and filters the results in memory

    Why it's wrong here

    Full table scans are avoided whenever possible in Delta Lake. Because the table is partitioned on the columns being filtered, the engine will leverage partition pruning to skip unnecessary data files entirely. Performing a full scan would be inefficient and negate the performance benefits of partitioning.

  • ✓

    It uses partition pruning to skip files in partitions not matching the filters

    Why this is correct

    The engine uses the filters on 'region' and 'order_date' to perform partition pruning. It only scans the storage directories corresponding to 'North' and dates from 2023 onwards. This significantly limits the volume of data read, resulting in much faster query execution and reduced I/O overhead.

  • ✗

    It re-indexes the table columns to facilitate faster lookup

    Why it's wrong here

    Databricks SQL does not use traditional database indexes in the way legacy RDBMS do. Instead, it uses file-level statistics and partition pruning to narrow down data access. Re-indexing is not an applicable concept in the context of Delta Lake, making this answer technically incorrect for Databricks.

  • ✗

    It automatically broadcasts the table to all nodes for faster processing

    Why it's wrong here

    Broadcasting is a join strategy used when one table is small enough to fit in memory on all nodes. It is not a mechanism for reading single tables from storage. Broadcasting does not apply to simple SELECT queries filtering from a single source table, regardless of partitioning.

About these practice questions

This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.