Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question
A data engineer has a PySpark DataFrame `readings` with columns `sensor_id` (string) and `celsius` (double). The engineer must produce a new DataFrame where every temperature is converted to Fahrenheit using the formula `celsius * 9/5 + 32`, while keeping both the original `sensor_id` and a column named `fahrenheit`, and must avoid collecting data to the driver. Which DataFrame operation should be used?
⚠ Common exam trap
The trap here is assuming that any transformation which computes a new value must use the RDD map() API, when the DataFrame column API already supports the same arithmetic lazily and with optimizer support.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
readings.select("sensor_id", (col("celsius") * 9/5 + 32).alias("fahrenheit"))
The requirement is a column-wise transformation that preserves sensor_id and adds fahrenheit, evaluated across executors. select() with a Column expression and alias is the canonical DataFrame API approach: it is lazy, Catalyst-optimized, and does not move data to the driver. Converting to the RDD API is unnecessary, dropping celsius changes the schema, and collect() pulls everything to the driver, which the scenario forbids.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
readings.collect().map(lambda r: (r.sensor_id, r.celsius * 9/5 + 32))
Why it's wrong here
collect() materializes every row on the driver, which is exactly what the scenario says to avoid and will fail or exhaust memory on a large DataFrame. The subsequent Python map() operates on a plain list, not a distributed DataFrame, so it does not produce a Spark DataFrame at all. This option violates both the distribution and schema requirements.
- ✗
readings.rdd.map(lambda r: (r.sensor_id, r.celsius * 9/5 + 32)).toDF(["sensor_id", "fahrenheit"])
Why it's wrong here
Converting to the RDD API discards the Catalyst optimizer, Tungsten code generation, and schema metadata that a DataFrame expression would retain. The map(lambda ...) also runs Python row-by-row, which is far slower than a native column expression, and toDF() must re-infer or restate the schema. It does avoid driver collection, but it is not the operation the scenario is asking for and loses DataFrame optimizations.
- ✓
readings.select("sensor_id", (col("celsius") * 9/5 + 32).alias("fahrenheit"))
Why this is correct
select() is a narrow transformation that projects existing columns and can also project arbitrary Column expressions with an alias, producing the requested schema without any driver-side collection. The arithmetic on the celsius Column is translated into a Catalyst expression and evaluated distributively on each executor partition, so the resulting DataFrame is lazy and only materializes when an action such as show() or write() is invoked.
- ✗
readings.withColumn("fahrenheit", col("celsius") * 9/5 + 32).drop("celsius")
Why it's wrong here
withColumn() would correctly add the fahrenheit column, but the scenario explicitly requires keeping the original sensor_id and producing fahrenheit; dropping celsius is not requested and changes the output schema. If the intent were only to add a column, withColumn() would be acceptable, but the drop() makes this option fail the stated requirement of preserving the input columns besides the derived one.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.