Databricks-Spark-Assoc Using Spark SQL Practice Question
An analyst notices that a Spark SQL query filtering a Delta table with `WHERE order_date = '2024-03-15'` scans far more data than expected, even though the table is partitioned by `order_date`. The partition column was loaded as a string in `yyyy-MM-dd` format. Which explanation best accounts for the excessive scan?
⚠ Common exam trap
The trap here is blaming a missing configuration flag for poor pruning, when pruning is automatic and fails only when the predicate and partition column types or formats are incompatible.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The literal is compared against a partition column of a different type or format, so pruning cannot match partitions.
Partition pruning depends on Catalyst proving that a filter predicate matches partition values, which requires the literal's type and format to align with the partition column. A string partition column compared to a differently typed or formatted literal prevents static pruning, causing a full scan. Correcting the literal's type and format lets the optimizer eliminate irrelevant partitions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The literal is compared against a partition column of a different type or format, so pruning cannot match partitions.
Why this is correct
If the partition column is stored as a string and the literal is cast or formatted differently, Catalyst may fail to derive a static partition predicate and fall back to scanning all partitions. Type or format mismatches between the filter literal and the partition column defeat pruning. Aligning the literal's type and format with the column restores efficient partition elimination.
- ✗
Partition pruning is disabled by default and must be enabled with `spark.sql.optimizer.enablePartitionPruning`.
Why it's wrong here
Partition pruning is applied automatically by the Catalyst optimizer when the filter references a partition column with a compatible literal; there is no such configuration flag to enable it. The real issue lies in how the predicate is expressed, not in a missing setting. This option misattributes the cause to a nonexistent toggle.
- ✗
The query lacks a `LIMIT` clause, which forces Spark to read every file in the table.
Why it's wrong here
`LIMIT` restricts the number of returned rows and does not control how many files or partitions are scanned for a filtered query. Adding a limit would not reduce the scan when the predicate itself cannot prune partitions. This option confuses result limiting with data skipping, which are unrelated mechanisms.
- ✗
Delta tables never support partition pruning; only Parquet tables do.
Why it's wrong here
Delta Lake is built on Parquet and fully supports partition pruning, along with data skipping via file statistics. Claiming Delta lacks pruning is factually wrong and would mislead the analyst away from the actual cause. The excessive scan stems from predicate compatibility, not from a Delta limitation.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.