Databricks-DE-Assoc Data Transformation and Modeling Practice Question
A data engineer is creating a Silver table in a Delta Live Tables pipeline. The pipeline must continuously ingest new files from a cloud storage location as they arrive, and the engineer wants to avoid reprocessing files that were already ingested. Which approach should be used to read the source data?
⚠ Common exam trap
The trap here is assuming that a simple batch read of a directory inside a streaming definition gives incremental file tracking.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Auto Loader with cloudFiles as the streaming source and let it track discovered files in its state store.
Auto Loader with cloudFiles is the purpose-built streaming source for incremental ingestion from cloud storage. It discovers new files, records processed files in a persistent state store so they are not read twice, and integrates with Delta Live Tables streaming tables. This provides continuous ingestion with exactly-once processing semantics and optional schema evolution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use Auto Loader with cloudFiles as the streaming source and let it track discovered files in its state store.
Why this is correct
Auto Loader is the recommended streaming source for cloud storage ingestion in Delta Live Tables. It incrementally discovers new files, persists a state store so already processed files are not reprocessed, and supports schema inference and evolution. This matches the requirement for continuous, incremental ingestion without duplicate processing.
- ✗
Use dbutils.fs.ls(path) in a loop and append new file paths to a Delta table before reading them.
Why it's wrong here
Manual directory listing with dbutils.fs.ls does not scale to large buckets and requires custom logic to detect which files are new. It also runs outside the pipeline's streaming semantics, so checkpointing and exactly-once behavior are lost. This approach is error-prone and reinvents functionality Auto Loader already provides.
- ✗
Define the source as a batch table and schedule the pipeline to run every minute with a full refresh.
Why it's wrong here
A full refresh reprocesses all source data on every run, which is expensive and defeats the goal of avoiding reprocessing. Batch tables in Delta Live Tables do not maintain incremental file state. Scheduling more frequently only increases cost and does not provide incremental ingestion semantics.
- ✗
Use spark.read.format("parquet").load(path) inside a streaming table definition.
Why it's wrong here
A batch read inside a streaming table definition is not a valid streaming source and will fail or force a full reprocessing pattern. Even if it ran, it would re-read all files on each refresh rather than tracking which files were already consumed. It does not provide incremental file discovery or state tracking.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.