Databricks-DE-Assoc Data Ingestion and Loading Practice Question
A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?
⚠ Common exam trap
The trap here is assuming that `CREATE TABLE AS SELECT` or `INSERT INTO` directly from a file path provides the same reliability as `COPY INTO`.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
`COPY INTO`
`COPY INTO` is the recommended command for loading files from cloud storage into Delta tables. It provides atomicity, idempotency, and file-level tracking, making it perfect for one-time or incremental batch ingestion. It also supports schema evolution and validation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
`COPY INTO`
Why this is correct
`COPY INTO` is designed for idempotent, atomic loads from cloud storage into Delta tables. It tracks previously loaded files and skips them on subsequent runs, making it ideal for one-time or incremental batch loads. It also provides schema validation and can handle large files efficiently.
- ✗
`CREATE TABLE AS SELECT` from the CSV path
Why it's wrong here
`CREATE TABLE AS SELECT` creates a new table and loads data in one step, but it is not idempotent and cannot be rerun without recreating the table. It does not track loaded files, so it is unsuitable for incremental or repeated loads. It also lacks schema evolution support.
- ✗
`INSERT INTO` with a `SELECT * FROM csv.` path ``
Why it's wrong here
`INSERT INTO` with a direct CSV path can work but is not atomic and lacks file-level tracking. It may lead to duplicates if rerun and does not provide the same level of reliability as `COPY INTO`. It is also less efficient for large files due to lack of optimizations.
- ✗
`MERGE INTO` using the CSV as source
Why it's wrong here
`MERGE INTO` is used for upserts between a source and target table, not for direct file ingestion. It requires the source data to be in a table or view. While it can be used after loading the CSV into a temporary table, it is not a direct ingestion command.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.