Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer runs a PySpark job on Databricks that reads a large Delta table, filters on a timestamp column, and writes results to another Delta table. The job takes 45 minutes, but the Spark UI shows that 90% of task time is spent reading from the source table. The developer wants to reduce the read time. Which action should the developer take?
⚠ Common exam trap
The trap here is assuming that caching or repartitioning will speed up a slow read, when the real solution is to reduce the amount of data read via data skipping.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Delta Lake data skipping by ensuring the timestamp column is in the table's partitioning or Z-ORDER BY columns.
The Spark UI indicates that the bottleneck is reading from the source Delta table. Delta Lake data skipping leverages file statistics to avoid reading irrelevant files when a filter predicate is applied. Ensuring the timestamp column is a partition column or has Z-ORDER BY applied allows the engine to prune files effectively, reducing I/O and overall job time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a higher value.
Why it's wrong here
The shuffle partitions setting controls the number of partitions created during wide transformations like joins or aggregations. The scenario describes a slow read phase, not a shuffle phase. Changing this parameter would not affect the time spent reading from the source table and could introduce unnecessary overhead elsewhere.
- ✓
Enable Delta Lake data skipping by ensuring the timestamp column is in the table's partitioning or Z-ORDER BY columns.
Why this is correct
Delta Lake data skipping uses file-level statistics to skip reading files that do not contain relevant data. If the timestamp column is a partition column or has been optimized with Z-ORDER BY, the query engine can prune files based on the filter predicate, drastically reducing I/O and read time. This directly addresses the observed bottleneck.
- ✗
Cache the source DataFrame in memory using .cache() before applying the filter.
Why it's wrong here
Caching the entire source DataFrame would require reading all data from storage first, which is the very step that is slow. Caching does not reduce the initial read cost and may cause memory pressure or spill. It benefits repeated access to the same data, not a single scan with a selective filter.
- ✗
Repartition the source DataFrame by the timestamp column before filtering.
Why it's wrong here
Repartitioning by the timestamp column would shuffle the entire dataset across the cluster, adding network and disk overhead. Since the filter is applied after reading, this does not reduce the volume of data read from the source. It may even increase total execution time and resource usage without addressing the root cause of slow reads.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.