Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

You are debugging a PySpark DataFrame job on Databricks that performs multiple transformations and actions on a large delta table. You notice that the execution plan shows redundant computations where the same upstream DataFrame is evaluated repeatedly. Which transformation should you apply to optimize this workflow and avoid recomputing the upstream lineage?

⚠ Common exam trap

Candidates often confuse caching with broadcasting or assume Spark automatically caches all reused DataFrames. Spark does not evaluate reference frequency and will recompute lineage unless explicitly told to persist.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call df.cache() or df.persist() before referencing the DataFrame in multiple downstream actions.

Persisting or caching a DataFrame instructs Spark to store intermediate results in memory or disk across actions, preventing expensive recomputation of the entire lineage graph. This optimization is critical for iterative algorithms or workflows where a single DataFrame is referenced multiple times downstream.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Call df.broadcast() on the DataFrame before each downstream join operation.

    Why it's wrong here

    Broadcasting sends a copy of a small table to all worker nodes to optimize joins by avoiding shuffle operations. It does not cache intermediate lineage results or prevent recomputation of expensive filter and map transformations across multiple separate actions.

  • ✗

    Call df.repartition() to distribute the data evenly across partitions before execution.

    Why it's wrong here

    Repartitioning forces a full shuffle across the cluster to change partition counts or key distributions. While useful for fixing data skew, it increases network overhead and does not store intermediate results in memory to avoid repeated lineage evaluation.

  • ✓

    Call df.cache() or df.persist() before referencing the DataFrame in multiple downstream actions.

    Why this is correct

    Caching serializes or stores the evaluated DataFrame partitions in memory or disk storage. When subsequent actions trigger execution, Spark retrieves the cached blocks directly instead of re-evaluating the entire upstream transformation DAG from source files.

  • ✗

    Call df.coalesce() to reduce the number of shuffle partitions prior to the final write operation.

    Why it's wrong here

    Coalesce only changes partition count before writing, so the upstream lineage still recomputes on each action. The stem needs caching to materialise the DataFrame once. Coalesce is the right call when reducing output files or shuffle partitions for a write, not for eliminating repeated evaluation of shared lineage.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.