Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer notices that a DataFrame transformation chain is executed twice: once for a count action used for logging and again for a write action. The source is a large Delta table and the repeated scan adds several minutes. Which action avoids the duplicate computation with the least risk?
⚠ Common exam trap
The trap here is trying to tune shuffle or join settings for a problem that is actually caused by recomputing the same lineage across two separate actions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call cache or persist on the transformed DataFrame before the count so the second action reuses the materialized data.
When the same DataFrame lineage feeds two actions, Spark recomputes it for each action unless the intermediate result is materialized. Calling cache or persist after the expensive transformations stores the result so the later write reads it directly, removing the second full scan of the Delta table. The change is local, reversible, and does not alter results, making it the lowest-risk option among those presented.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the DataFrame to a Pandas DataFrame for the count and then write from the original Spark DataFrame.
Why it's wrong here
Collecting to Pandas pulls the entire dataset to the driver, which will fail or exhaust driver memory on a large Delta table. It also does not prevent the write action from recomputing the lineage. Mixing Pandas and Spark here adds serialization cost and risk without removing the duplicate scan. This is a dangerous and ineffective approach.
- ✓
Call cache or persist on the transformed DataFrame before the count so the second action reuses the materialized data.
Why this is correct
Caching the transformed DataFrame materializes it once, so the subsequent write action reads the cached partitions instead of recomputing the entire lineage from the Delta table. This directly eliminates the duplicate scan. It is a targeted change with predictable memory cost and no change to results. For a DataFrame reused across multiple actions, caching is the standard remedy.
- ✗
Set spark.sql.shuffle.partitions equal to the number of executor cores to speed up each execution.
Why it's wrong here
Partition count tuning changes parallelism within a single job execution; it does not prevent the second action from recomputing the same transformations. The problem is duplication across actions, not slow tasks within one action. Lowering partitions to the core count can also underutilize the cluster. This leaves the redundant scan in place.
- ✗
Increase spark.sql.autoBroadcastJoinThreshold so more joins are broadcast and the plan becomes cheaper.
Why it's wrong here
Broadcast thresholds affect join strategy selection, not whether a DataFrame's lineage is recomputed across actions. The duplicate work here comes from two actions triggering two separate executions of the same lineage. Raising the threshold could even cause out-of-memory errors if a large table is broadcast. It does not address the repeated scan.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.