Databricks-Spark-Assoc Pandas API on Spark Practice Question
A developer is using the pandas API on Spark in a Databricks notebook. They have a pandas-on-Spark DataFrame psdf with a default index. They call psdf.sort_values('amount') and then psdf.head(10). They observe that the resulting index values are not sequential from 0 to 9, but instead appear as arbitrary integers. What is the most likely explanation for this behavior?
⚠ Common exam trap
The trap here is assuming that pandas-on-Spark mimics pandas' default behavior of producing a sequential index after sorting, when the default 'distributed' index intentionally preserves arbitrary but stable integers.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The default index type is 'distributed', which assigns arbitrary but stable integers to rows and does not guarantee sequential values after operations like sort_values.
The default index type in pandas API on Spark is 'distributed', which assigns stable but non-sequential integers to rows. When you sort a DataFrame, the existing index values travel with the rows, so the resulting index is not 0..n-1. To get sequential indices after sorting, you must explicitly call reset_index() or set the index type to 'distributed-sequence' before sorting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The default index type is 'distributed-sequence', which preserves the original row order and does not reassign new sequential indices after sorting.
Why it's wrong here
The 'distributed-sequence' index type does generate sequential indices, but it is not the default. The default is 'distributed' when the index is not set, which assigns non-sequential, arbitrary integers that are stable across operations but not contiguous. With 'distributed-sequence', you would see contiguous values, but it requires an extra shuffle and is not enabled by default.
- ✗
The default index type is 'sequence', which creates a single-partition sequence index that is lost during shuffles such as sorting.
Why it's wrong here
The 'sequence' index type is not a valid option in pandas API on Spark. The available types are 'distributed', 'distributed-sequence', and 'distributed-sequence'. The 'sequence' type does not exist. Additionally, sort_values does not lose the index; it simply carries the existing index values with the sorted rows, which for the default 'distributed' type are non-sequential.
- ✗
The default index type is 'distributed-sequence', but sort_values triggers a shuffle that resets the index to arbitrary values because sequence indices cannot survive shuffles.
Why it's wrong here
The default is not 'distributed-sequence'; it is 'distributed'. Even if it were 'distributed-sequence', sequence indices are designed to survive shuffles by recomputing the sequence, though with additional overhead. The observed behavior of arbitrary integers is characteristic of the 'distributed' index type, which assigns stable but non-sequential integers. The premise about sequence indices resetting is incorrect.
- ✓
The default index type is 'distributed', which assigns arbitrary but stable integers to rows and does not guarantee sequential values after operations like sort_values.
Why this is correct
By default, pandas-on-Spark uses the 'distributed' index type when no index is specified. This type assigns each row a unique integer that is stable across operations but not necessarily sequential or contiguous. After sort_values, the index values are carried along with the rows, so they appear out of order. This is expected behavior and differs from pandas, where sorting resets the index only if reset_index is called.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.