Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer has a DataFrame `df` with columns `order_id`, `customer_id`, and `order_total`. They need to create a new DataFrame containing only the rows where `order_total` is greater than 100 and only the columns `order_id` and `order_total`. Which combination of DataFrame operations accomplishes this most efficiently?
⚠ Common exam trap
The trap here is mixing up filter with select ordering or referencing a column after it has been projected away, which produces an AnalysisException rather than the intended subset.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
df.filter(df.order_total > 100).select("order_id", "order_total")
Chaining filter (or where) with select on the DataFrame API keeps both transformations narrow, enables Catalyst to push the predicate down to the source, and prunes unneeded columns. The alternative using an aggregation introduces a shuffle, the RDD variant abandons Catalyst optimizations, and the remaining option filters on the wrong column and references a dropped field.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
df.where("order_total > 100").groupBy("order_id", "order_total").count()
Why it's wrong here
Introducing groupBy forces a wide shuffle and changes the output schema by adding a count column, which the scenario does not request. This is dramatically less efficient than a simple filter-and-project chain and produces semantically different results, since aggregation is unnecessary for a pure row-and-column selection.
- ✗
df.select("order_id", "order_total").where(df.customer_id > 100)
Why it's wrong here
This applies the numeric condition to customer_id instead of order_total, so the wrong rows are kept. Additionally, referencing df.customer_id after the select attempts to use a column that no longer exists in the projected DataFrame, which raises an AnalysisException at action time.
- ✗
df.rdd.filter(lambda r: r.order_total > 100).map(lambda r: (r.order_id, r.order_total)).toDF()
Why it's wrong here
Dropping to the RDD API bypasses Catalyst and Tungsten, losing predicate pushdown, column pruning, and whole-stage code generation. It is also more verbose and error-prone, since attribute access on Row objects is fragile. The DataFrame-native filter and select chain is strictly preferable for this simple projection-and-filter task.
- ✓
df.filter(df.order_total > 100).select("order_id", "order_total")
Why this is correct
filter followed by select is the idiomatic, lazy transformation chain. Catalyst can push the filter below the projection and even down to the data source when the format supports predicate pushdown, minimizing I/O. Both operations are narrow transformations, so no shuffle occurs, and the resulting DataFrame contains exactly the requested rows and columns.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.