Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A developer is building a feature that must, for each `customer_id`, concatenate the distinct `product` values from many rows into a single comma-separated string. The result must contain each product only once per customer. Which single approach produces this result?
⚠ Common exam trap
The trap here is reaching for collect_list out of habit, forgetting that it preserves duplicates, whereas the distinct requirement demands collect_set.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
df.groupBy("customer_id").agg(concat_ws(",", collect_set("product")).alias("products"))
Producing a distinct comma-separated list per customer requires an aggregation that deduplicates first. collect_set returns a set of unique product values, which concat_ws then joins into one string. collect_list keeps duplicates, and the other approaches either collapse to a single value or reshape the data into columns, none of which yield the required distinct concatenated string.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
df.groupBy("customer_id").agg(concat_ws(",", collect_set("product")).alias("products"))
Why this is correct
collect_set gathers distinct product values into an array per customer, and concat_ws joins that array into a comma-separated string. Because collect_set deduplicates, each product appears once per customer, exactly matching the requirement to build a distinct list in a single aggregated column.
- ✗
df.groupBy("customer_id").agg(concat_ws(",", collect_list("product")).alias("products"))
Why it's wrong here
collect_list preserves duplicates, so a product purchased many times would appear many times in the concatenated string. The scenario explicitly requires each product only once per customer, so this produces a valid but incorrect result with repeated entries and inflates the string length.
- ✗
df.groupBy("customer_id").pivot("product").agg(count("product"))
Why it's wrong here
pivot turns distinct product values into separate columns with counts, producing a wide table rather than a single comma-separated string column. The output shape and content are entirely different from the requested per-customer concatenated list, so it does not satisfy the requirement.
- ✗
df.select("customer_id", concat_ws(",", "product")).groupBy("customer_id").agg(first("product"))
Why it's wrong here
concat_ws applied to a single string column does not aggregate across rows; it operates row by row. Taking first afterward returns only one product per customer, discarding the rest, so the concatenated list is never built and the result loses information.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.