Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A developer has two DataFrames, `orders` (columns `order_id`, `customer_id`) and `customers` (columns `customer_id`, `customer_name`). They need a result containing every order, with the matching customer name where one exists and null where the customer is not found. Which single join configuration guarantees this?
⚠ Common exam trap
The trap here is treating an inner join as harmless because most orders do match, when even a single unmatched order is silently dropped and the requirement explicitly demands every order be kept.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
orders.join(customers, on="customer_id", how="left")
Preserving all rows of the left DataFrame while enriching with matching values from the right is the definition of a left join. Unmatched orders retain their fields and receive nulls for customer_name. Inner, right, and full joins either drop unmatched orders or introduce extra customer-only rows, so they do not meet the requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
orders.join(customers, on="customer_id", how="left")
Why this is correct
A left join keeps every row from orders regardless of whether a match exists in customers. When no matching customer_id is found, the customer_name column is populated with null, precisely the behavior the scenario requires. The on parameter also avoids a duplicate customer_id column in the result.
- ✗
orders.join(customers, on="customer_id", how="right")
Why it's wrong here
A right join preserves all rows from customers, not orders. Orders referencing a missing customer would be excluded, and customers with no orders would appear with null order columns, which inverts the intended direction and does not satisfy the requirement to retain every order.
- ✗
orders.join(customers, on="customer_id", how="inner")
Why it's wrong here
An inner join returns only orders that have a matching customer. Orders without a corresponding customer record would be silently dropped, violating the requirement to keep every order. This is the most common mistake when the intent is to preserve the left DataFrame fully.
- ✗
orders.join(customers, on="customer_id", how="full")
Why it's wrong here
A full outer join keeps unmatched rows from both sides, so customers with no orders would also appear, introducing rows the scenario does not want. It preserves orders, but the extra customer-only rows change the result set and the row count, so it is not the correct choice here.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.