Courseiva

Databricks-Spark-Assoc Developing DataFrame/DataSet API Applications Practice Question

A developer has a DataFrame `raw` and needs to permanently persist it as Parquet partitioned by `region`, overwriting any existing data at that path, without registering it in the metastore. Which call achieves this?

⚠ Common exam trap

Watch out — candidates often confuse `saveAsTable`, which registers metastore metadata, with a direct path-based `write`, which does not.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

raw.write.mode("overwrite").partitionBy("region").parquet("/data/out")

The DataFrameWriter path with `mode("overwrite")`, `partitionBy("region")`, and `.parquet` writes typed Parquet files into per-region directories and replaces prior contents in one step. Because it targets a filesystem path rather than a table name, no metastore entry is created, satisfying the "without registering it in the metastore" constraint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    raw.rdd.saveAsTextFile("/data/out")

    Why it's wrong here

    `saveAsTextFile` operates on the underlying RDD and writes plain text with a string representation of each row, discarding schema and column types. It also cannot express partitioning by `region` or Parquet encoding, so downstream readers lose the typed columns and the directory layout entirely.

  • ✗

    raw.write.format("parquet").saveAsTable("out")

    Why it's wrong here

    `saveAsTable` registers a managed or external table in the metastore, which the developer explicitly wants to avoid, and it omits both partitioning and the overwrite mode. Without `partitionBy`, no region directory structure is produced, and without overwrite the write fails when the target already exists.

  • ✗

    raw.createOrReplaceTempView("out").write.parquet("/data/out")

    Why it's wrong here

    `createOrReplaceTempView` returns `None`, so chaining `.write` on it raises an `AttributeError` at runtime. The view is also session-scoped metadata only, not a persistence mechanism, and this snippet never applies partitioning or an overwrite mode, so it fails on every count in the scenario.

  • ✓

    raw.write.mode("overwrite").partitionBy("region").parquet("/data/out")

    Why this is correct

    This uses the DataFrameWriter with `mode("overwrite")` to replace existing data, `partitionBy("region")` to lay out Hive-style directories, and the `parquet` format sink to write files. It writes directly to the filesystem path without touching the metastore, which is precisely the requirement stated in the scenario.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.