Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer is building a PySpark job that reads a Parquet dataset with 2,000 files into `df`. They call `df.cache()` and then execute three separate actions in the same session. The Spark UI shows the Parquet files are read from storage three times, and the cache never appears in the Storage tab. What is the most likely cause?

⚠ Common exam trap

The trap here is assuming `cache()` is a property of the data itself, when it is actually attached to a specific logical plan node.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

`cache()` was applied to `df` but a variable reassignment such as `df = df.withColumn(...)` created a new logical plan that no longer references the cached node.

Caching in Spark marks a specific logical plan node as persistent. Any subsequent transformation that derives a new DataFrame from it produces a different plan that does not include the cached node, so actions against the new DataFrame bypass the cache entirely. The tell-tale signs here are repeated Parquet reads and an empty Storage tab. To fix this, either cache after the final transformation or reuse the exact same cached DataFrame reference across actions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    `cache()` was applied to `df` but a variable reassignment such as `df = df.withColumn(...)` created a new logical plan that no longer references the cached node.

    Why this is correct

    This is correct because `cache()` marks a specific logical plan node. If `df` is later reassigned to a transformation like `withColumn`, the new DataFrame references an uncached plan, so actions re-read Parquet. The original cached node is never triggered, which explains both the repeated file reads and the empty Storage tab.

  • ✗

    `cache()` only takes effect after `persist(StorageLevel.MEMORY_ONLY)` is called, so the default call does nothing.

    Why it's wrong here

    This is incorrect because `cache()` is shorthand for `persist(StorageLevel.MEMORY_AND_DISK)` and is fully effective on its own. No additional `persist` call is required. The behavior described in the scenario is not caused by the absence of a second persist call but by the cached plan node never being executed.

  • ✗

    `cache()` stores data in memory only, and the dataset is too large, so Spark silently drops the cached blocks.

    Why it's wrong here

    Spark's `cache()` defaults to MEMORY_AND_DISK, so blocks that do not fit in memory spill to disk rather than being dropped. Even if memory were insufficient, the Storage tab would still show cached partitions with disk usage. The absence of any cache entry indicates the cache was never materialized, not that blocks were evicted.

  • ✗

    The Parquet files are stored in a format that Spark cannot cache, so `cache()` is a no-op for Parquet sources.

    Why it's wrong here

    This is false: Spark can cache any DataFrame regardless of source format, including Parquet. The storage level applies to the in-memory columnar representation of the DataFrame, not to the underlying file format. Parquet is fully cacheable and commonly cached in production pipelines for repeated access.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.