A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?
Filtering right after the read lets Spark push the predicate into the Delta scan where possible and reduces the number of rows that must be shuffled for the window. This lowers shuffle size, memory pressure, and runtime, and it is the direct way to get the filter to act before the window partitioning.
Why this answer
The physical plan shows the filter after the window, meaning the window shuffle processes all rows. Moving the filter to immediately after the read lets Spark push it into the Delta scan where possible and shrink the dataset before the window shuffle, reducing shuffle volume, memory use, and overall runtime.
Exam trap
The trap here is tuning window spill or repartitioning by the window key, when the real win is applying the filter before the window so fewer rows ever enter the shuffle.