Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer must combine two DataFrames, `left_df` and `right_df`, on a key column `id`. They need every row from `left_df` regardless of whether a match exists in `right_df`, and matching rows from `right_df` where available, with unmatched right-side columns filled as null. Which join invocation produces this result?

⚠ Common exam trap

The trap here is assuming that left_outer and full_outer are interchangeable when only one side's unmatched rows are relevant, which quietly adds or removes rows from the result.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

left_df.join(right_df, on="id", how="left_outer")

A left outer join retains every left row and fills right-side columns with null when no match exists, which is exactly what the scenario specifies. Inner join discards unmatched left rows, right outer join preserves the wrong side, and full outer join introduces unmatched right rows that were not requested.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    left_df.join(right_df, on="id", how="right_outer")

    Why it's wrong here

    A right outer join preserves all rows from right_df, which is the opposite of the requirement. Rows in left_df without a match would be dropped, and unmatched right rows would be retained with null left columns, inverting the desired behavior.

  • ✗

    left_df.join(right_df, on="id", how="inner")

    Why it's wrong here

    An inner join returns only rows where the key exists in both DataFrames, discarding any left rows without a match. This contradicts the requirement to retain all rows from left_df, so the result would silently drop unmatched left records.

  • ✓

    left_df.join(right_df, on="id", how="left_outer")

    Why this is correct

    A left outer join preserves every row from left_df, attaching matching right-side columns where the key matches and filling them with null otherwise. This exactly satisfies the stated requirement that all left rows survive while unmatched right columns become null.

  • ✗

    left_df.join(right_df, on="id", how="full_outer")

    Why it's wrong here

    A full outer join keeps unmatched rows from both sides, so right_df rows with no counterpart in left_df would also appear. The scenario only requires preserving left rows and matching right rows, so full_outer introduces extra rows that are not wanted.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.