Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer is writing a PySpark job that must read a Parquet dataset, apply several transformations, and then write the result back to storage. They want to ensure the schema of the written data is inferred directly from the DataFrame rather than from any external definition, and they want to append to an existing Parquet directory. Which write configuration accomplishes this?

⚠ Common exam trap

The trap here is assuming `mergeSchema` is a write option, when it is actually applied when reading Parquet files with differing schemas.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.write.mode('append').format('parquet').save('/path/output')

Appending Parquet output with the DataFrame writer uses the DataFrame's own schema because Parquet stores schema metadata in each file. The overwrite mode would destroy existing data, JSON output changes the storage format, and `mergeSchema` is a read-side option that does not affect how the DataFrame schema is written. The append-mode Parquet write is the only configuration that meets both conditions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    df.write.mode('append').format('json').save('/path/output')

    Why it's wrong here

    Writing in JSON format does not preserve a typed schema in the same way Parquet does; JSON is text-based and schema is inferred at read time from the data itself. The scenario specifically involves a Parquet dataset and expects Parquet output, so switching formats changes storage characteristics and loses Parquet's columnar benefits. The mode is correct, but the format is not.

  • ✗

    df.write.mode('append').option('mergeSchema', 'true').parquet('/path/output')

    Why it's wrong here

    `mergeSchema` is a read-time option for reconciling differing Parquet schemas across files, not a write-time setting that changes how the DataFrame schema is applied. On write, the DataFrame schema is already used directly. Including this option at write time is ineffective and may mislead developers into thinking schema handling is being customized when it is not.

  • ✗

    df.write.mode('overwrite').format('parquet').save('/path/output')

    Why it's wrong here

    `overwrite` mode deletes the existing directory contents before writing, so it does not append to prior data. Although the schema is still derived from the DataFrame, the destructive behavior violates the append requirement. This mode is appropriate when a full refresh is intended, not when incremental additions are needed.

  • ✓

    df.write.mode('append').format('parquet').save('/path/output')

    Why this is correct

    Writing with the DataFrame writer in `append` mode and the `parquet` format stores the DataFrame using its own schema, since Parquet embeds the schema in the file metadata. Appending adds new files to the existing directory without removing prior data. This directly satisfies both requirements: schema derived from the DataFrame and append semantics.

About these practice questions

This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.