Databricks-Spark-Assoc Using Spark SQL Practice Question
A data scientist is using Spark SQL to compute a running total of sales amounts for each customer, ordered by transaction date. The query must return, for each row, the sum of all previous sales for that customer up to and including the current row. Which window specification should be used?
⚠ Common exam trap
A common mix-up: candidates confuse ROWS with RANGE, or using a frame that sums future rows instead of preceding ones, which yields incorrect cumulative totals when duplicate dates exist.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
OVER (PARTITION BY customer_id ORDER BY transaction_date ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW)
A running total requires a window frame that starts at the beginning of the partition and ends at the current row. The specification with ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW, combined with PARTITION BY customer_id and ORDER BY transaction_date, achieves this by summing all rows up to and including the current one for each customer. Other frames either sum future rows, include tied rows, or limit to a small window, none of which produce the desired cumulative sum.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
OVER (PARTITION BY customer_id ORDER BY transaction_date ROWS BETWEEN CURRENT ROW AND UNBOUNDED FOLLOWING)
Why it's wrong here
This frame starts at the current row and extends to the end of the partition, computing a reverse cumulative sum (sum of current and future rows). It does not produce a running total of previous sales. The data scientist needs the sum of all prior rows including the current one, so this frame is the opposite of what is required.
- ✓
OVER (PARTITION BY customer_id ORDER BY transaction_date ROWS BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW)
Why this is correct
This window specification partitions by customer_id, orders by transaction_date, and defines the frame from the start of the partition up to the current row. SUM(sales_amount) over this window yields a running total per customer. The ROWS frame is explicit and ensures each row includes all preceding rows, which is exactly the requirement. It is the standard approach for cumulative sums.
- ✗
OVER (PARTITION BY customer_id ORDER BY transaction_date ROWS BETWEEN 1 PRECEDING AND CURRENT ROW)
Why it's wrong here
This frame includes only the immediately preceding row and the current row, computing a two-row moving sum. It does not accumulate all previous sales. The requirement is a running total from the beginning, so a frame of one preceding row is insufficient. It would produce incorrect results for customers with more than two transactions.
- ✗
OVER (PARTITION BY customer_id ORDER BY transaction_date RANGE BETWEEN UNBOUNDED PRECEDING AND CURRENT ROW)
Why it's wrong here
RANGE with UNBOUNDED PRECEDING and CURRENT ROW includes all rows with the same transaction_date as the current row, not just the physical rows up to the current one. If multiple transactions share the same date, they are all included, which may overcount. For a precise running total per row, ROWS is preferred. RANGE can be correct in some cases, but it does not guarantee row-by-row accumulation when ties exist.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.