Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A developer has a DataFrame `events` with columns `user_id`, `event_type`, and `payload` (a JSON string). They need to extract the `device` field from `payload` into a new column without changing the other columns or the row count. Which approach is correct?
⚠ Common exam trap
The trap here is reaching for `from_json` and `explode` reflexively, when a single-field extraction with `get_json_object` is simpler and preserves the DataFrame shape.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
events.withColumn("device", get_json_object("payload", "$.device"))
Extracting a single scalar from a JSON string column is exactly what `get_json_object` does, and `withColumn` adds the result without altering other columns or row counts. Parsing the whole payload with `from_json` is unnecessary when only one field is needed, and any use of `explode` or RDD conversion changes the shape of the DataFrame.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
events.withColumn("device", explode(from_json("payload", schema)))
Why it's wrong here
`explode` is for arrays or maps and would change the row count by generating one row per element. Applying it to a struct from `from_json` is not valid and would not produce a single scalar device value. This option violates the requirement to preserve the row count and the original columns.
- ✗
events.select("user_id", "event_type", from_json("payload", schema).getField("device"))
Why it's wrong here
`from_json` returns a struct column, and calling `.getField("device")` on the column is not valid PySpark API; the correct function is `getField` used differently or `col("parsed.device")`. This option also drops the original `payload` column and does not assign a name to the new field, so the result would not preserve the original schema plus a new column as required.
- ✗
events.rdd.map(lambda r: r.payload["device"]).toDF()
Why it's wrong here
`r.payload` is a string, not a dict, so subscripting it with `["device"]` would raise a TypeError at runtime. Even if it worked, this converts to RDD and back to DataFrame, losing the original schema and columns, and it defeats Catalyst optimization. It does not meet the preserve-columns and preserve-rows requirement.
- ✓
events.withColumn("device", get_json_object("payload", "$.device"))
Why this is correct
`get_json_object` extracts a scalar value from a JSON string using a JSONPath expression. Wrapping it in `withColumn` adds the `device` column while preserving all existing columns and the row count. This is the simplest correct approach when only a single field is needed from the JSON string without parsing the entire document.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.