Courseiva
Pandas API on Spark →hardMultiple Choice

Databricks-Spark-Assoc Pandas API on Spark Practice Question

You are working with a Pandas-on-Spark DataFrame `psdf` that has a default index generated by Spark. You need to perform a join with another Pandas-on-Spark DataFrame `other` that also has a default index. After the join, you notice that the resulting DataFrame has a new index and the original indices are lost. Which of the following best explains this behavior?

⚠ Common exam trap

The trap here is assuming that the default index behaves like a pandas index and is preserved across joins, but it is not stable across shuffles.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The default index is not preserved across joins because it is not a true pandas index; it is a synthetic index generated per partition, and joins require shuffling which discards the original index.

The default index in Pandas API on Spark is a synthetic index that is not preserved across operations that shuffle data, such as joins. When you perform a join, the data is redistributed across partitions, and the original default index is discarded. The resulting DataFrame gets a new default index. To maintain a stable index, you should set an explicit index using `set_index` before the join.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The default index is lost because the join operation uses the index as the join key by default, and since both DataFrames have the same default index, it causes a conflict.

    Why it's wrong here

    This is incorrect because the join operation does not automatically use the index as the join key. In Pandas API on Spark, you must explicitly specify the join columns using `on` parameter. The default index is not used as a join key unless specified. The loss of index is due to shuffling, not a conflict.

  • ✓

    The default index is not preserved across joins because it is not a true pandas index; it is a synthetic index generated per partition, and joins require shuffling which discards the original index.

    Why this is correct

    This is correct. In Pandas API on Spark, when no explicit index is set, a default index is created using `distributed-sequence` or `distributed` index types. These indices are not stable across operations that require shuffling, such as joins. The join operation shuffles data based on join keys, and the default index is not carried over, resulting in a new default index for the output.

  • ✗

    The default index is preserved across joins only if you set the Spark configuration `spark.pandas.join.index` to true.

    Why it's wrong here

    This is incorrect because there is no such configuration `spark.pandas.join.index` in Pandas API on Spark. While there are configurations to control index behavior, such as `spark.pandas.default.index.type`, there is no option to preserve the default index across joins. The default index is inherently unstable across shuffles.

  • ✗

    The default index is lost because the join operation converts both DataFrames to Spark DataFrames, performs the join, and then converts back to Pandas-on-Spark, which resets the index.

    Why it's wrong here

    This is incorrect because while the join is executed on the underlying Spark DataFrames, the conversion back to Pandas-on-Spark does not inherently reset the index. The index is lost because the default index is not preserved through the shuffle. The internal conversion is an implementation detail, but the root cause is the shuffling of data.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.