Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer has a DataFrame `df` with columns `id` and `score`, and several rows contain null values in `score`. They want a new DataFrame in which rows with a null `score` are removed, keeping only rows where `score` is present. Which single call achieves this?

⚠ Common exam trap

The trap here is mixing up DataFrame.drop, which removes columns, with na.drop, which removes rows, because both share the word drop and are easy to confuse under time pressure.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.na.drop(subset=["score"])

Dropping rows based on nulls in a specific column is done with na.drop and the subset argument, which limits the null check to the named columns. Filling values keeps the rows, filtering for nulls keeps the wrong rows, and dropping a column removes the field entirely, so none of those match the stated goal of discarding incomplete rows.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    df.na.fill(0, subset=["score"])

    Why it's wrong here

    na.fill replaces nulls with the supplied value instead of removing rows. The scenario asks for rows with null score to be dropped, not substituted, so this changes the data semantics and keeps every row. Downstream aggregates would then be skewed by the artificial zeros.

  • ✗

    df.filter(col("score").isNull())

    Why it's wrong here

    This filter keeps only rows where score is null, the exact opposite of what is required. The developer wants to retain non-null scores, so applying this would discard all valid records and leave only the incomplete ones, producing a nearly empty DataFrame.

  • ✗

    df.drop("score")

    Why it's wrong here

    drop with a string argument removes the entire column named score from the schema rather than removing rows. The result would have no score column at all, so any later computation referencing score fails with an unresolved column error, which does not satisfy the requirement.

  • ✓

    df.na.drop(subset=["score"])

    Why this is correct

    na.drop with the subset argument removes rows where any of the listed columns contain null, so restricting it to score drops exactly the rows missing a score while leaving other columns untouched. It returns a new DataFrame and does not mutate the original, matching the requirement precisely.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.