Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?
⚠ Common exam trap
The trap here is assuming that enabling an optimizer feature automatically reuses results across actions, when only an explicit cache or persist call stores computed partitions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call `df.persist(StorageLevel.MEMORY_AND_DISK)` before the loop and `df.unpersist()` after it.
Because the same DataFrame is acted on repeatedly and the underlying Delta table does not change, persisting the DataFrame keeps the computed partitions available across iterations, so the lineage is evaluated once rather than 40 times. Repartitioning, enabling adaptive execution, or round-tripping through RDDs all leave the recomputation behavior intact.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Wrap the DataFrame in `spark.createDataFrame(df.rdd)` so the derived object keeps the computed rows in the driver.
Why it's wrong here
Converting to an RDD and back does not cache anything; it simply rewrites the lineage and may add serialization overhead. The new DataFrame would still recompute from the source on every action, and pulling results toward the driver would risk memory pressure for a large dataset.
- ✓
Call `df.persist(StorageLevel.MEMORY_AND_DISK)` before the loop and `df.unpersist()` after it.
Why this is correct
Caching the DataFrame before the loop stores the computed partitions in executor memory (spilling to disk when needed), so each subsequent action reads the cached blocks instead of walking the lineage again. Since the source is unchanged, the cached result stays valid for all 40 iterations, and unpersisting releases the memory once the loop finishes.
- ✗
Set `spark.sql.adaptive.enabled` to true so the optimizer reuses prior query results automatically.
Why it's wrong here
Adaptive Query Execution re-optimizes physical plans at runtime based on shuffle statistics, but it does not persist or reuse results across separate actions. Each `count()` and `show()` still launches a new job that recomputes the lineage from the Delta table, so the repeated work in the loop is not eliminated.
- ✗
Add `.repartition(200)` to the DataFrame before the loop so each iteration reads a different partition set.
Why it's wrong here
Repartitioning only changes how rows are distributed across partitions; it does not store any computed result. Every action inside the loop would still trigger a full shuffle and then recompute the entire upstream lineage, so the repeated execution cost remains and is actually made worse by the added exchange.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.