Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer has a PySpark DataFrame `events` with columns `user_id`, `event_time` (timestamp), and `payload` (string). They must produce a new DataFrame where each row is enriched with the `payload` value from the user's immediately preceding event, ordered by `event_time`, without collapsing any rows. Which TWO approaches accomplish this? (Choose two.)
⚠ Common exam trap
The trap here is assuming that any 'previous row' lookup works globally, when Window functions must be partitioned by the entity whose sequence you care about.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a Window partitioned by `user_id` and ordered by `event_time`, then use `lag('payload', 1).over(windowSpec)` inside `withColumn`.
Enriching rows with a prior value without collapsing requires a Window function. Partitioning by `user_id` and ordering by `event_time` ensures the 'previous' row is the same user's immediately earlier event, while `lag` returns the payload from that row and preserves row count. Both `lag(...).over(windowSpec)` and the inline `Window.partitionBy(...).orderBy(...)` expression produce identical semantics; the alternatives either collapse rows, use invalid syntax, or lose per-user ordering.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use `df.orderBy('event_time').rdd.zipWithIndex()` and manually look up the previous index per user in a driver-side dictionary.
Why it's wrong here
This fails because `zipWithIndex` assigns a global index after a global sort, not a per-user sequence, so the 'previous index' does not correspond to the user's prior event. Additionally, collecting a driver-side dictionary scales poorly and defeats the purpose of distributed Spark processing for a large `events` table.
- ✓
Create a Window partitioned by `user_id` and ordered by `event_time`, then use `lag('payload', 1).over(windowSpec)` inside `withColumn`.
Why this is correct
This is correct because a Window partitioned by `user_id` and ordered by `event_time` restricts `lag` to each user's own timeline, and `lag('payload', 1)` returns the prior row's payload while preserving every row. `withColumn` adds the result as a new column without aggregating, which satisfies the requirement of not collapsing rows.
- ✗
Apply `window('user_id', 'event_time')` with `lag('payload')` and rely on Spark to fill nulls for the first event.
Why it's wrong here
This is invalid because `window` is not a function that accepts column names; the correct API is `Window.partitionBy('user_id').orderBy('event_time')`. Also, the frame specification matters: default lag still returns null for the first row, but the syntax as written would not even compile, so it cannot satisfy the scenario.
- ✓
Use `df.withColumn('prev_payload', lag('payload').over(Window.partitionBy('user_id').orderBy('event_time')))`.
Why this is correct
This is correct because it constructs the Window directly inside `over`, partitioned by `user_id` and ordered by `event_time`, which scopes `lag` to each user's ordered events. The result is added as a new column while retaining all original rows, exactly matching the requirement to enrich without collapsing.
- ✗
Call `groupBy('user_id').agg(collect_list('payload'))` and then explode the resulting list back into rows.
Why it's wrong here
This fails because `groupBy` plus `collect_list` collapses rows per user and loses the original row-level ordering unless you also sort inside the aggregation, which is not guaranteed. Exploding the list also loses the association with the original `event_time` values, so the enriched DataFrame would not correctly align payloads to their preceding events.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.