Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?

⚠ Common exam trap

The trap here is tuning window spill or repartitioning by the window key, when the real win is applying the filter before the window so fewer rows ever enter the shuffle.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply the filter on the timestamp column immediately after reading the Delta table, before calling the window function, so predicate pushdown and early filtering reduce rows entering the shuffle.

The physical plan shows the filter after the window, meaning the window shuffle processes all rows. Moving the filter to immediately after the read lets Spark push it into the Delta scan where possible and shrink the dataset before the window shuffle, reducing shuffle volume, memory use, and overall runtime.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase spark.sql.windowExec.buffer.spill.threshold so the window operator can spill more data to disk.

    Why it's wrong here

    This setting controls spilling behavior within the window operator but does not change when the filter is applied. The shuffle still processes the full unfiltered dataset, so the dominant cost remains, and tuning spill thresholds only mitigates memory pressure rather than reducing input volume.

  • ✗

    Repartition the DataFrame by customer_id before applying the filter so the window avoids a shuffle.

    Why it's wrong here

    Repartitioning by the window key can align data for the window, but it shuffles the full unfiltered dataset. Filtering first reduces the rows that participate in any shuffle, so performing the repartition before the filter is strictly more expensive and does not achieve the goal of early filtering.

  • ✓

    Apply the filter on the timestamp column immediately after reading the Delta table, before calling the window function, so predicate pushdown and early filtering reduce rows entering the shuffle.

    Why this is correct

    Filtering right after the read lets Spark push the predicate into the Delta scan where possible and reduces the number of rows that must be shuffled for the window. This lowers shuffle size, memory pressure, and runtime, and it is the direct way to get the filter to act before the window partitioning.

  • ✗

    Cache the DataFrame before the window and call count() to materialize it, then apply the filter.

    Why it's wrong here

    Caching materializes the full unfiltered dataset and adds an action, increasing memory and time without moving the filter earlier in the plan. The window still shuffles all rows, so this adds overhead rather than reducing the data entering the shuffle.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.