Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer notices that a PySpark DataFrame transformation chain runs a full scan of a Delta table each time a new action is invoked, even though the source data has not changed. The developer wants to persist the intermediate DataFrame in memory across actions. Which method should be used?
⚠ Common exam trap
The trap here is conflating checkpointing, which truncates lineage to disk, with caching, which keeps the materialized DataFrame available in memory for repeated actions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call DataFrame.cache() before the first action so the DataFrame is stored in memory on first computation.
cache() persists the DataFrame across actions using the default MEMORY_AND_DISK level, so the first action computes and stores partitions and later actions read from the cache rather than rescanning the Delta table. Checkpointing writes to disk and cuts lineage but does not keep data in memory, while AQE and collect() do not provide reusable in-memory persistence.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Call DataFrame.collect() after each transformation to materialize the results on the driver.
Why it's wrong here
collect() returns all rows to the driver as a local collection, which can exhaust driver memory on large datasets and does not register the data as a cached DataFrame for later Spark actions. It is a terminal action, not a persistence mechanism, so it would not allow subsequent DataFrame actions to reuse the intermediate result.
- ✓
Call DataFrame.cache() before the first action so the DataFrame is stored in memory on first computation.
Why this is correct
cache() marks the DataFrame for in-memory persistence with the default MEMORY_AND_DISK storage level, so the first action materializes it and subsequent actions reuse the cached partitions instead of rescanning the Delta table. This directly addresses the repeated full scans described in the scenario and is the idiomatic way to persist across multiple actions.
- ✗
Set spark.sql.adaptive.enabled to true so the optimizer reuses prior scan results automatically.
Why it's wrong here
Adaptive Query Execution optimizes the physical plan at runtime, such as coalescing shuffle partitions and choosing join strategies, but it does not cache scan results across separate actions. Each action still triggers a new scan of the Delta table. This setting does not provide cross-action data reuse.
- ✗
Call DataFrame.checkpoint() to truncate the lineage and write the data to a reliable file system.
Why it's wrong here
checkpoint() writes the DataFrame to a reliable storage path and cuts lineage, which is useful for long lineages or iterative algorithms, but it does not keep the data in memory for fast reuse. Subsequent actions read from the checkpoint files rather than memory, so it does not provide the in-memory persistence the scenario requires.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.