Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A data engineer has a DataFrame `events` with columns `user_id` and `ts`. They need to add a column `prev_ts` that holds the previous event timestamp for each user, ordered by `ts` ascending, without collapsing rows. Which operation accomplishes this?

⚠ Common exam trap

The trap here is reaching for `sortWithinPartitions` or `lead` and assuming partition-local ordering plus the wrong offset function reproduces the previous-value semantics of a properly windowed `lag`.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Window.partitionBy('user_id').orderBy('ts') with lag('ts', 1) via withColumn

Using a window partitioned by `user_id` and ordered by `ts` with the `lag` function returns the previous event's timestamp for each row while retaining all rows. The other choices either aggregate away detail, rely on partition-local sorting that cannot guarantee correct ordering, or drop rows entirely. Only the windowed `lag` approach yields a per-row previous timestamp at the original granularity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    events.groupBy('user_id').agg(max('ts').alias('prev_ts'))

    Why it's wrong here

    Aggregation collapses each user's rows into a single output row, so the original per-event granularity is lost. It also computes the maximum timestamp rather than the preceding one for each event. Since the requirement explicitly forbids collapsing rows, this operation cannot satisfy the scenario and would discard needed detail.

  • ✓

    Window.partitionBy('user_id').orderBy('ts') with lag('ts', 1) via withColumn

    Why this is correct

    A window partitioned by `user_id` and ordered by `ts`, combined with the `lag` function, returns the previous row's timestamp while preserving every original row. This is exactly the pattern for adding a prior-value column without collapsing data, and it executes as a distributed window operation under Catalyst.

  • ✗

    events.sortWithinPartitions('user_id', 'ts').withColumn('prev_ts', lead('ts'))

    Why it's wrong here

    Sorting within partitions only orders data inside each existing partition and does not guarantee that all rows for a user land in the same partition or in a defined sequence. `lead` returns the next value, not the previous one. Together these produce incorrect or nondeterministic results for the prior-timestamp requirement.

  • ✗

    events.dropDuplicates(['user_id']).withColumnRenamed('ts', 'prev_ts')

    Why it's wrong here

    Dropping duplicates by `user_id` keeps only one row per user, which is the opposite of preserving all events. Renaming the timestamp column does not compute any previous value at all. This approach destroys data and produces a meaningless `prev_ts`, so it fails the requirement completely.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.