Courseiva
Importing Data →hardMultiple Choice

Databricks-DA-Assoc Importing Data Practice Question

An analyst is loading Parquet files from cloud storage into a Delta table using COPY INTO. The source directory contains files with several different schemas because upstream teams added columns over time. The analyst wants the table to accept the union of all columns without manual intervention. Which COPY INTO behavior should the analyst configure?

⚠ Common exam trap

It's easy for candidates to confuse overwrite behavior with schema merging, when overwrite replaces data and mergeSchema extends the table definition.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set mergeSchema to true in the COPY INTO options so new columns are added to the target table.

The mergeSchema option in COPY INTO adds any columns found in the source files to the target table, producing the union of columns across heterogeneous files. This matches the requirement to accept new columns without manual ALTER TABLE statements, while retaining all previously loaded data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable overwrite mode so each run replaces the table with the latest file's schema.

    Why it's wrong here

    Overwrite mode replaces the table contents on each run, which would discard previously loaded data and leave only the schema of the most recent batch. That is the opposite of accepting the union of all columns across files. The analyst needs accumulation, not replacement, so this option would lose history and still not merge schemas.

  • ✗

    Create the target table with a manually defined superset schema containing every possible column.

    Why it's wrong here

    Hand-defining a superset schema requires knowing every column in advance, which contradicts the goal of avoiding manual intervention. If upstream adds another column later, the table would again reject the new files. This approach also forces the analyst to maintain the schema definition as a separate artifact, adding ongoing maintenance burden.

  • ✗

    Use a permissive mode that silently drops columns not present in the target table schema.

    Why it's wrong here

    Dropping unknown columns would let the load succeed but would discard the new data, which defeats the purpose of capturing the union of columns. The analyst wants the extra columns retained in the table, not ignored. Silent dropping also hides schema drift from the team, making it harder to detect upstream changes.

  • ✓

    Set mergeSchema to true in the COPY INTO options so new columns are added to the target table.

    Why this is correct

    COPY INTO supports a mergeSchema option that automatically adds columns present in the source files but missing from the target table. This is exactly the union-of-columns behavior the analyst wants, and it avoids manual ALTER TABLE steps. Without it, files containing extra columns would be rejected or would fail schema matching.

About these practice questions

Courseiva writes every Databricks-DA-Assoc question from scratch — 291 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.