Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer is building a feature pipeline and needs to add a monotonically increasing integer column `row_num` to a DataFrame `df` that assigns consecutive numbers to rows within each partition of a specified ordering, similar to a window function. Which approach uses the DataFrame API to compute this value?
⚠ Common exam trap
The trap here is treating monotonically_increasing_id() as a drop-in row counter, when its values are only increasing and unique rather than consecutive, and it cannot honor an ordering clause.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
df.withColumn("row_num", row_number().over(Window.orderBy("some_col")))
Consecutive numbering based on a defined ordering is the job of the row_number() window function applied with over(Window.orderBy(...)). It yields one-based consecutive integers according to the window specification, which is what the feature pipeline needs. monotonically_increasing_id() leaves gaps and ignores ordering, a literal constant does not increment, and the RDD zipWithIndex approach collapses to one partition and bypasses DataFrame optimizations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
df.withColumn("row_num", monotonically_increasing_id())
Why it's wrong here
monotonically_increasing_id() generates unique, increasing 64-bit integers, but the values are not guaranteed to be consecutive: they encode the partition index in the upper bits, so gaps appear between partitions. The scenario asks for consecutive numbering within a specified ordering, which this function does not honor because it does not accept an ordering at all. It also does not restart per group.
- ✓
df.withColumn("row_num", row_number().over(Window.orderBy("some_col")))
Why this is correct
row_number() is a window function that assigns consecutive integers starting at one according to the window's ordering, which matches the requested behavior. Used with over(Window.orderBy(...)), it produces the running sequence the engineer needs. Note that a global orderBy in the window moves all rows to a single partition, so the engineer should be aware of that cost, but the semantics are exactly correct.
- ✗
df.withColumn("row_num", lit(1).cast("int"))
Why it's wrong here
This assigns the constant 1 to every row, which is not a sequence at all. It does not increment and does not depend on any ordering, so it fails the requirement of consecutive numbering. It is included here as a contrast to the window function approach, and it would only be useful if a literal placeholder column were needed.
- ✗
df.rdd.zipWithIndex().toDF()
Why it's wrong here
zipWithIndex() on an RDD does produce consecutive indices, but it requires a single partition internally, forcing all data through one executor and losing DataFrame optimizer benefits. Converting back with toDF() also loses the original schema unless tuples are reshaped. It is a heavyweight RDD workaround rather than the DataFrame API operation the scenario asks for, and it does not respect a specified ordering.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.