Databricks-DA-Assoc Importing Data Practice Question
A data analyst is setting up a recurring ingestion job that must load only new files from a cloud storage directory into a Delta table, must track which files have already been processed, and must fail clearly if the source schema changes unexpectedly. Which TWO capabilities should the analyst rely on? (Choose two.)
⚠ Common exam trap
The trap here is reaching for a manual notebook or full-table rebuild when the managed ingestion tools already provide file tracking and schema policy handling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Auto Loader with a checkpoint location tracks processed files and supports schema enforcement or evolution policies.
Both COPY INTO and Auto Loader maintain ingestion state so only new files are processed, and both can be configured to surface unexpected schema changes rather than silently accepting them. COPY INTO uses table-level load history, while Auto Loader uses a checkpoint directory with an explicit schema policy. Either satisfies the incremental and schema-safety requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Mounting the storage location and using spark.read on the directory with overwrite mode.
Why it's wrong here
Reading the whole directory with overwrite mode reprocesses every file on each run and discards prior data, so it provides neither incremental tracking nor a schema-change safeguard. It also bypasses Delta's load history. The scenario needs only new files to be loaded and schema drift to surface as a failure, which this approach cannot guarantee.
- ✗
CREATE TABLE AS SELECT on the source directory rebuilds the table from all files on every run.
Why it's wrong here
CREATE TABLE AS SELECT reads the entire source each time and replaces the table contents, so it does not track processed files and would reprocess everything. It also offers no built-in schema-change failure behavior beyond whatever the query itself encounters. This approach is wasteful and does not satisfy the incremental requirement.
- ✗
A scheduled notebook that lists the directory and compares filenames to a manually maintained table.
Why it's wrong here
Hand-rolling file tracking in a notebook duplicates functionality that COPY INTO and Auto Loader provide natively, and it is error-prone: concurrent runs, partial failures, and path normalization can all corrupt the manual ledger. The scenario asks for reliable tracking and clear schema failure, which managed ingestion tools deliver without custom code.
- ✓
Auto Loader with a checkpoint location tracks processed files and supports schema enforcement or evolution policies.
Why this is correct
Auto Loader uses a checkpoint directory to record which files it has consumed, enabling incremental processing of new arrivals. Its schema handling can be configured to fail on unexpected changes or to evolve, which matches the requirement to fail clearly when the source schema shifts. This makes it suitable for the recurring job described.
- ✓
COPY INTO maintains a load history so previously ingested file paths are skipped on subsequent runs.
Why this is correct
COPY INTO records ingested file paths in the table's transaction log and skips them on later executions. This provides exactly the incremental behavior the scenario requires: each run processes only files it has not seen before, without the analyst writing custom bookkeeping. It is the idempotent foundation for recurring file-based loads into Delta.
About these practice questions
This Databricks-DA-Assoc question is part of Courseiva's 291-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DA-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DA-Assoc exam.