Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer has a DataFrame `events` with columns `user_id` and `ts`. They need to add a column `prev_ts` that holds the previous event timestamp for each user, ordered by `ts` ascending, without collapsing rows. Which operation accomplishes this?
⚠ Common exam trap
The trap here is reaching for `sortWithinPartitions` or `lead` and assuming partition-local ordering plus the wrong offset function reproduces the previous-value semantics of a properly windowed `lag`.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Window.partitionBy('user_id').orderBy('ts') with lag('ts', 1) via withColumn
Using a window partitioned by `user_id` and ordered by `ts` with the `lag` function returns the previous event's timestamp for each row while retaining all rows. The other choices either aggregate away detail, rely on partition-local sorting that cannot guarantee correct ordering, or drop rows entirely. Only the windowed `lag` approach yields a per-row previous timestamp at the original granularity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
events.groupBy('user_id').agg(max('ts').alias('prev_ts'))
Why it's wrong here
Aggregation collapses each user's rows into a single output row, so the original per-event granularity is lost. It also computes the maximum timestamp rather than the preceding one for each event. Since the requirement explicitly forbids collapsing rows, this operation cannot satisfy the scenario and would discard needed detail.
- ✓
Window.partitionBy('user_id').orderBy('ts') with lag('ts', 1) via withColumn
Why this is correct
A window partitioned by `user_id` and ordered by `ts`, combined with the `lag` function, returns the previous row's timestamp while preserving every original row. This is exactly the pattern for adding a prior-value column without collapsing data, and it executes as a distributed window operation under Catalyst.
- ✗
events.sortWithinPartitions('user_id', 'ts').withColumn('prev_ts', lead('ts'))
Why it's wrong here
Sorting within partitions only orders data inside each existing partition and does not guarantee that all rows for a user land in the same partition or in a defined sequence. `lead` returns the next value, not the previous one. Together these produce incorrect or nondeterministic results for the prior-timestamp requirement.
- ✗
events.dropDuplicates(['user_id']).withColumnRenamed('ts', 'prev_ts')
Why it's wrong here
Dropping duplicates by `user_id` keeps only one row per user, which is the opposite of preserving all events. Renaming the timestamp column does not compute any previous value at all. This approach destroys data and produces a meaningless `prev_ts`, so it fails the requirement completely.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.