Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?

⚠ Common exam trap

The trap here is assuming that enabling an optimizer feature automatically reuses results across actions, when only an explicit cache or persist call stores computed partitions.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call `df.persist(StorageLevel.MEMORY_AND_DISK)` before the loop and `df.unpersist()` after it.

Because the same DataFrame is acted on repeatedly and the underlying Delta table does not change, persisting the DataFrame keeps the computed partitions available across iterations, so the lineage is evaluated once rather than 40 times. Repartitioning, enabling adaptive execution, or round-tripping through RDDs all leave the recomputation behavior intact.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Wrap the DataFrame in `spark.createDataFrame(df.rdd)` so the derived object keeps the computed rows in the driver.

    Why it's wrong here

    Converting to an RDD and back does not cache anything; it simply rewrites the lineage and may add serialization overhead. The new DataFrame would still recompute from the source on every action, and pulling results toward the driver would risk memory pressure for a large dataset.

  • ✓

    Call `df.persist(StorageLevel.MEMORY_AND_DISK)` before the loop and `df.unpersist()` after it.

    Why this is correct

    Caching the DataFrame before the loop stores the computed partitions in executor memory (spilling to disk when needed), so each subsequent action reads the cached blocks instead of walking the lineage again. Since the source is unchanged, the cached result stays valid for all 40 iterations, and unpersisting releases the memory once the loop finishes.

  • ✗

    Set `spark.sql.adaptive.enabled` to true so the optimizer reuses prior query results automatically.

    Why it's wrong here

    Adaptive Query Execution re-optimizes physical plans at runtime based on shuffle statistics, but it does not persist or reuse results across separate actions. Each `count()` and `show()` still launches a new job that recomputes the lineage from the Delta table, so the repeated work in the loop is not eliminated.

  • ✗

    Add `.repartition(200)` to the DataFrame before the loop so each iteration reads a different partition set.

    Why it's wrong here

    Repartitioning only changes how rows are distributed across partitions; it does not store any computed result. Every action inside the loop would still trigger a full shuffle and then recompute the entire upstream lineage, so the repeated execution cost remains and is actually made worse by the added exchange.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.