Be able to choose and write the correct DataFrame or Dataset transformation for filtering, windowing, and joining. The most important thing is matching the method and its arguments to the exact result: null handling, window partitioning and ordering, and the right join type.
Start practicing
Developing DataFrame/DataSet API Applications — choose a session length
Free · No account required
Domain overview
This domain covers building transformations with Spark DataFrames and Datasets on Databricks: filtering nulls, window functions, broadcast joins, and join types. Questions present a scenario and ask which PySpark DataFrame or Dataset API call produces the required result, so you must recognize the correct method and its arguments.
Exam objectives
Filtering null rows using df.filter(col('score').isNotNull()) or dropna(subset=['score'])
Adding previous-row values with Window.partitionBy('user_id').orderBy('ts') and lag()
Optimizing small-table joins with broadcast() hints or broadcast join thresholds
Selecting semi and anti joins via left_semi and left_anti join types
Using dropna() without subset drops rows with nulls in any column, not just the intended score column.
Forgetting to import pyspark.sql.functions and Window, or omitting partitionBy, which makes lag() span all users.
Assuming a normal join filters unmatched keys; left_semi keeps only matches and left_anti keeps only non-matches.
Click any question to see the full explanation and answer options, or start a focused practice session above.
You have a large DataFrame containing user transaction logs. You need to read this data and immediately repartition it by user_id to optimize downstream filtering operations. Which DataFrame API method should you use?
2You need to add a new calculated column 'discounted_price' to an existing DataFrame 'df' by multiplying 'price' by 0.9. Which DataFrame transformation accomplishes this correctly?
3A data engineer needs to join two large DataFrames, `sales` and `products`, on the `product_id` column. The `products` DataFrame is extremely small and fits entirely in a single executor's memory. To optimize performance and avoid a costly shuffle join across the network, which strategy should be applied using the Spark DataFrame API?
4You are processing a large, highly skewed PySpark DataFrame in Databricks and want to optimize a forthcoming join operation against a small lookup dimension table. Which TWO strategies are valid and effective DataFrame API techniques to optimize this join performance? (Choose TWO)
5A data engineer has a PySpark DataFrame `events` with columns `user_id`, `event_time` (timestamp), and `payload` (string). They must produce a new DataFrame where each row is enriched with the `payload` value from the user's immediately preceding event, ordered by `event_time`, without collapsing any rows. Which TWO approaches accomplish this? (Choose two.)
6A developer has a PySpark DataFrame `df` with columns `order_id`, `customer_id`, and `order_ts` (timestamp). They need to return only the most recent order per customer, keeping all original columns, and they want to avoid a self-join or a manual sort-then-dropDuplicates approach. Which DataFrame operation should they use?
7A developer is building a PySpark job that reads a Parquet dataset with 2,000 files into `df`. They call `df.cache()` and then execute three separate actions in the same session. The Spark UI shows the Parquet files are read from storage three times, and the cache never appears in the Storage tab. What is the most likely cause?
8A developer has a DataFrame `events` with columns `user_id`, `event_type`, and `payload` (a JSON string). They need to extract the `device` field from `payload` into a new column without changing the other columns or the row count. Which approach is correct?
9A data engineer has a PySpark DataFrame `orders` with a string column `order_ts` formatted as `yyyy-MM-dd HH:mm:ss`. They need a new column `order_date` containing only the date portion as a `date` type, so downstream code can filter by day. Which approach is correct?
10A data engineer has a PySpark DataFrame `events` with columns `event_ts` (TimestampType) and `user_id`. They must produce a new DataFrame where each row shows the event and the timestamp of that same user's previous event, ordered by `event_ts` within each `user_id`. Which code snippet correctly accomplishes this?
11A data engineer has a PySpark DataFrame `readings` with columns `sensor_id` (string) and `celsius` (double). The engineer must produce a new DataFrame where every temperature is converted to Fahrenheit using the formula `celsius * 9/5 + 32`, while keeping both the original `sensor_id` and a column named `fahrenheit`, and must avoid collecting data to the driver. Which DataFrame operation should be used?
12A developer must join a 4 TB `transactions` DataFrame against a 900 MB `merchants` DataFrame on `merchant_id`. The cluster has 40 executors each with 16 GB of memory, and the job currently shuffles the large side. They want to avoid the shuffle entirely. Which change should they make?
13A developer has a DataFrame `raw` and needs to permanently persist it as Parquet partitioned by `region`, overwriting any existing data at that path, without registering it in the metastore. Which call achieves this?
14A developer needs to join `orders` (large) with `customers` (small, a few thousand rows) on `customer_id`. They want to broadcast the small table and confirm the broadcast actually took effect. Which combination of actions is correct?
15A developer needs to add a column `full_name` to DataFrame `people` by concatenating `first_name` and `last_name` with a single space, and must handle rows where `last_name` is null by producing just the first name. Which expression is correct?
16A data engineer has a DataFrame `events` with columns `user_id` and `ts`. They need to add a column `prev_ts` that holds the previous event timestamp for each user, ordered by `ts` ascending, without collapsing rows. Which operation accomplishes this?
17A data engineer has a batch DataFrame `df` with a `status` string column and wants to keep only rows where `status` equals "active". The engineer wants the filter applied as early as possible in the plan and does not want a shuffle. Which operation best fits this requirement?
18A data engineer is writing a PySpark job that must read Parquet data, apply transformations, and then write the result. They want to avoid recomputation if the resulting DataFrame is referenced multiple times in later stages. Which action best accomplishes this?
19A data engineer has a DataFrame `orders` with columns `order_id`, `customer_id`, and `amount`, and a small lookup DataFrame `tiers` with `customer_id` and `tier`. The engineer wants to attach the tier to every order. Some orders have a customer_id that is not present in tiers, and those orders must still appear with a null tier. Which operation produces this result?
20A data engineer has two DataFrames: `orders` with columns `order_id` and `customer_id`, and `customers` with columns `customer_id` and `region`. They must produce a result containing only orders whose `customer_id` exists in `customers`, keeping every matching order row exactly once with no customer columns added. Which operation should be used?
21A data engineer is building a feature pipeline and needs to add a monotonically increasing integer column `row_num` to a DataFrame `df` that assigns consecutive numbers to rows within each partition of a specified ordering, similar to a window function. Which approach uses the DataFrame API to compute this value?
22A developer has a DataFrame `df` with columns `id` and `score`, and several rows contain null values in `score`. They want a new DataFrame in which rows with a null `score` are removed, keeping only rows where `score` is present. Which single call achieves this?
23A developer is building a feature that must, for each `customer_id`, concatenate the distinct `product` values from many rows into a single comma-separated string. The result must contain each product only once per customer. Which single approach produces this result?
24A data engineer has a DataFrame `df` with columns `order_id`, `customer_id`, and `order_total`. They need to create a new DataFrame containing only the rows where `order_total` is greater than 100 and only the columns `order_id` and `order_total`. Which combination of DataFrame operations accomplishes this most efficiently?
25A developer has two DataFrames, `orders` (columns `order_id`, `customer_id`) and `customers` (columns `customer_id`, `customer_name`). They need a result containing every order, with the matching customer name where one exists and null where the customer is not found. Which single join configuration guarantees this?
26A developer has a DataFrame `trades` with columns `trade_id`, `symbol`, and `price`. They want to add a column `prev_price` containing the price of the immediately preceding trade for the same symbol, ordered by `trade_id` ascending, without collapsing rows. Which transformation should they use?
27A developer must combine two DataFrames, `left_df` and `right_df`, on a key column `id`. They need every row from `left_df` regardless of whether a match exists in `right_df`, and matching rows from `right_df` where available, with unmatched right-side columns filled as null. Which join invocation produces this result?
28A developer has a DataFrame `df` and wants to remove duplicate rows considering only the columns `user_id` and `event_type`, keeping the first occurrence according to the current row order. Which DataFrame operation achieves this?
29A data engineer needs to add a column `rank_in_dept` to a DataFrame `employees` that ranks each employee by `salary` descending within their `department`, but they must not collapse rows. They also want ties to receive the same rank with gaps afterward. Which expression correctly produces this column using the DataFrame API?
30A developer is writing a PySpark job that must read a Parquet dataset, apply several transformations, and then write the result back to storage. They want to ensure the schema of the written data is inferred directly from the DataFrame rather than from any external definition, and they want to append to an existing Parquet directory. Which write configuration accomplishes this?
31A data engineer is building a PySpark application that must validate incoming records in a DataFrame `raw` before loading them into a curated table. They want to apply user-defined validation logic that cannot be expressed with built-in functions, and they want the result to remain a DataFrame column of Boolean values. Which TWO approaches allow this? (Choose two.)
Be able to choose and write the correct DataFrame or Dataset transformation for filtering, windowing, and joining. The most important thing is matching the method and its arguments to the exact result: null handling, window partitioning and ordering, and the right join type.
The Courseiva Databricks-Spark-Assoc question bank contains 31 questions in the Developing DataFrame/DataSet API Applications domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Developing DataFrame/DataSet API Applications domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included