Databricks-Spark-Assoc Using Spark SQL Practice Question
A developer has a Delta table events with a high-cardinality column user_id and a low-cardinality column country. A query filters on country = 'US' and also on user_id IN (...). The developer runs EXPLAIN and sees a full scan of all files. Which statement about data skipping and the Delta table's statistics correctly explains why the filter on country is not skipping files?
⚠ Common exam trap
The trap here is assuming that a low-cardinality filter column automatically triggers file skipping, when skipping actually depends on the presence and usefulness of per-file statistics in the transaction log.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Delta data skipping uses per-file min/max statistics, and if the table was written without collecting statistics or the file count is small, the country filter cannot eliminate files even though the value is low cardinality.
Delta data skipping prunes files using per-file statistics such as min/max and null counts recorded in the transaction log. A filter on country can only eliminate files if those statistics were collected at write time and the file layout separates country values. Low cardinality does not by itself enable skipping, and skipping is not limited to join conditions or partition columns, nor does it require a special enablement flag for basic operation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Delta data skipping only works on columns used in JOIN conditions, so a WHERE filter on country is never eligible for file pruning regardless of statistics.
Why it's wrong here
Data skipping applies to predicate filters in WHERE clauses as well as join conditions, provided the relevant statistics are available in the transaction log. There is no restriction that limits skipping to join keys. The real issue in this scenario is the availability and usefulness of statistics, not the clause type where the predicate appears.
- ✗
Data skipping requires the filtered column to be the first column in the table's partitioning scheme, so country must be a partition column for any file pruning to occur.
Why it's wrong here
Partitioning is a separate mechanism from file-level statistics. Data skipping can prune files using min/max statistics on non-partition columns, so country does not need to be the first partition column. Partitioning is one way to organize data, but it is not a prerequisite for statistic-based skipping, which works from the transaction log metadata.
- ✗
Data skipping is disabled by default and must be enabled with spark.databricks.delta.dataSkipping.enabled before any file pruning can happen on any column.
Why it's wrong here
Data skipping is a built-in Delta capability that operates from the statistics stored in the transaction log and does not require a configuration flag to be turned on for basic pruning. The absence of pruning here stems from missing or ineffective statistics, not from a disabled feature. Referencing a non-existent enablement setting misdiagnoses the cause.
- ✓
Delta data skipping uses per-file min/max statistics, and if the table was written without collecting statistics or the file count is small, the country filter cannot eliminate files even though the value is low cardinality.
Why this is correct
Delta data skipping relies on per-file statistics such as min/max and null counts stored in the transaction log. If statistics were not collected, or if there are too few files for skipping to matter, the optimizer cannot prune files for country = 'US'. Low cardinality alone does not guarantee skipping; the statistics must exist and the file layout must be granular enough to separate values.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.