Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer runs a PySpark job on Databricks that filters a Delta table and then calls count() and show(). The Spark UI shows two separate scans of the same table. The developer wants to avoid the second scan without changing the filter logic. Which action should be taken?
⚠ Common exam trap
The trap here is assuming that adaptive execution or read-partition tuning will prevent a second scan, when only an explicit persist or cache keeps the intermediate DataFrame available across actions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call persist() with StorageLevel.MEMORY_AND_DISK on the filtered DataFrame before count() and show().
Persisting the filtered DataFrame with MEMORY_AND_DISK materializes it on the first action and lets the second action read the stored partitions instead of rescanning the Delta table. AQE, maxPartitionBytes, and checkpointing do not provide in-memory reuse across actions; they change plan shape or write to disk rather than caching the intermediate result for repeated access.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Call checkpoint() on the filtered DataFrame to write it to a reliable path before count().
Why it's wrong here
checkpoint() writes the DataFrame to storage and truncates lineage, which can help with long lineages but does not keep data in memory for fast reuse. Subsequent actions read from the checkpoint files on disk, so while the second scan of the Delta table is avoided, the performance benefit is disk-based and the intent of caching in memory is not achieved.
- ✗
Enable spark.sql.adaptive.enabled and set spark.sql.adaptive.coalescePartitions.enabled to true.
Why it's wrong here
AQE coalesces shuffle partitions and optimizes join strategies at runtime, but it does not cache scan results across separate actions. Each action still plans and executes its own scan of the Delta table, so the duplicate scan would remain. This setting addresses post-shuffle partition sizing, not cross-action reuse.
- ✓
Call persist() with StorageLevel.MEMORY_AND_DISK on the filtered DataFrame before count() and show().
Why this is correct
persist() with MEMORY_AND_DISK materializes the filtered DataFrame on the first action and reuses the stored partitions for the second action, eliminating the duplicate scan. It is equivalent in effect to cache() but with an explicit storage level, and it directly targets the repeated scan shown in the Spark UI without altering the filter logic.
- ✗
Set spark.sql.files.maxPartitionBytes to 256 MB so each scan reads fewer, larger partitions.
Why it's wrong here
maxPartitionBytes controls how input files are split into partitions, affecting the number and size of read tasks. It does not prevent a second action from rescanning the table. Even with larger partitions, count() and show() would each trigger their own scan, so the duplicate read observed in the UI would persist.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.