Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A data engineer is writing a PySpark job that must read Parquet data, apply transformations, and then write the result. They want to avoid recomputation if the resulting DataFrame is referenced multiple times in later stages. Which action best accomplishes this?

⚠ Common exam trap

The trap here is believing that triggering an action such as `count()` persists the data, when in fact only `cache` or `persist` retains the computed partitions for reuse.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Call df.cache() or df.persist() before the DataFrame is reused

Marking the DataFrame with `cache` or `persist` stores its computed partitions so subsequent actions reuse them rather than replaying the full lineage. Forcing a count, converting to an RDD, or checkpointing without a configured directory either fails to store reusable data or introduces errors and overhead. Caching is the standard mechanism for avoiding recomputation across multiple downstream uses.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Call df.checkpoint() without setting a checkpoint directory

    Why it's wrong here

    `checkpoint` truncates lineage by writing to reliable storage, but it requires a checkpoint directory to be configured, and calling it without one raises an error. Even when configured, checkpointing is heavier than caching and is intended for lineage truncation rather than repeated reuse. It does not reliably serve the stated purpose here.

  • ✗

    Call df.count() to force computation before reuse

    Why it's wrong here

    Calling `count()` triggers a full job but does not store the result anywhere reusable; the next action recomputes the lineage from scratch. It only confirms the row count and wastes a full pass over the data. This does not prevent recomputation and therefore fails the stated objective.

  • ✗

    Call df.rdd to convert the DataFrame to an RDD

    Why it's wrong here

    Converting to an RDD does not cache anything and typically bypasses Catalyst optimizations, potentially making later operations slower. The RDD is still lazily evaluated, so reuse triggers recomputation of the underlying plan. This conversion does not achieve the caching goal and introduces unnecessary overhead.

  • ✓

    Call df.cache() or df.persist() before the DataFrame is reused

    Why this is correct

    Caching or persisting materializes the DataFrame's computed partitions so later actions reuse them instead of re-executing the full lineage. This directly addresses the goal of avoiding recomputation when the same DataFrame feeds multiple downstream operations, and it remains lazy until an action triggers the actual caching.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.