Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer needs to add a column `rank_in_dept` to a DataFrame `employees` that ranks each employee by `salary` descending within their `department`, but they must not collapse rows. They also want ties to receive the same rank with gaps afterward. Which expression correctly produces this column using the DataFrame API?
⚠ Common exam trap
The trap here is treating `rank` and `dense_rank` as interchangeable, when only `rank` leaves gaps after tied values.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
employees.withColumn('rank_in_dept', rank().over(Window.partitionBy('department').orderBy(col('salary').desc())))
The requirement calls for a window function that ranks within each department, orders by descending salary, keeps all rows, and leaves gaps after ties. `rank()` combined with a window partitioned by department and ordered by descending salary satisfies every condition. `dense_rank` omits gaps, `row_number` breaks ties arbitrarily, and aggregating with a window function is not valid.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
employees.withColumn('rank_in_dept', row_number().over(Window.partitionBy('department').orderBy(col('salary').desc())))
Why it's wrong here
`row_number()` assigns a unique sequential number to every row, even when salaries are tied, so tied employees receive different ranks. This violates the requirement that ties share the same rank. It is useful when a strict total ordering is needed, but it does not implement the tie-aware ranking described in the scenario.
- ✓
employees.withColumn('rank_in_dept', rank().over(Window.partitionBy('department').orderBy(col('salary').desc())))
Why this is correct
`rank()` assigns the same rank to tied values and leaves gaps in the sequence after ties, which matches the requirement. Using `Window.partitionBy('department').orderBy(col('salary').desc())` scopes the ranking per department and orders by descending salary. `withColumn` preserves all original rows, so no data is collapsed. This is the precise combination for the scenario.
- ✗
employees.withColumn('rank_in_dept', dense_rank().over(Window.partitionBy('department').orderBy(col('salary').desc())))
Why it's wrong here
`dense_rank()` assigns the same rank to ties but does not leave gaps, so after a tie the next rank is consecutive rather than skipped. The scenario explicitly requires gaps after ties, which `dense_rank` does not produce. It is otherwise a valid window function, but it does not satisfy the stated tie-handling rule.
- ✗
employees.groupBy('department').agg(rank().over(Window.orderBy(col('salary').desc())).alias('rank_in_dept'))
Why it's wrong here
Window functions cannot be used inside `agg()` in this manner, and this expression would also collapse rows per department, losing individual employee records. The window definition lacks partitioning by department, so ranking would be global. This is both syntactically invalid for the intended purpose and semantically wrong for the requirement.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.