Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?
⚠ Common exam trap
Candidates frequently confuse dataframe lineage optimization or broadcast hints with caching, missing that reused DataFrames need explicit persistence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the cache() or persist() method on the DataFrame.
Caching (or persisting) the DataFrame keeps it in memory or on disk for subsequent actions. This is essential for iterative algorithms or workloads where the same data is reused, as it avoids recomputing the lineage from scratch. Understanding the trade-offs between memory and disk persistence is fundamental for building performant Databricks pipelines that maximize resource reuse and minimize redundant computation time.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of cores per executor.
Why it's wrong here
Increasing cores per executor improves parallelism for a single task but does not address redundant computations. If the data is recomputed multiple times, the job will remain slow regardless of core counts because the underlying DAG is being re-evaluated repeatedly, wasting time on identical transformations across multiple steps.
- ✓
Use the cache() or persist() method on the DataFrame.
Why this is correct
Caching pins the DataFrame in memory or disk, allowing subsequent actions to read the computed data directly rather than re-running the entire lineage. This significantly improves performance for iterative processing tasks, reducing latency and avoiding repeated data reads and transformations that occur in non-cached, re-evaluated Spark DataFrames.
- ✗
Switch the file format from Parquet to CSV.
Why it's wrong here
CSV is a text-based, row-oriented format that is much slower than Parquet for Spark. Parquet supports column pruning and predicate pushdown, which are critical for performance. Moving to CSV would actually decrease performance by removing these optimizations and increasing the amount of I/O required for every operation.
- ✗
Enable speculative execution.
Why it's wrong here
Speculative execution is designed to mitigate the impact of straggler tasks in a cluster with uneven performance. It does not help with performance in scenarios where the same data is processed repeatedly. It is a cluster-level fault tolerance mechanism, not a data-caching optimization for iterative operations.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.