Databricks-DE-Assoc Data Ingestion and Loading Practice Question
A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?
⚠ Common exam trap
The trap here is assuming that increasing maxFilesPerTrigger speeds up ingestion; it actually controls batch size, not discovery.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable cloudFiles.useIncrementalListing.
Incremental listing leverages cloud storage APIs to list only new files, drastically reducing the time to discover files in large directories. Other options affect processing or schema handling but do not optimize file discovery, which is the bottleneck in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable cloudFiles.useIncrementalListing.
Why this is correct
This option enables incremental listing, which uses the cloud storage's native file listing capabilities to discover only new files since the last run. It significantly improves performance when dealing with large numbers of files because it avoids full directory scans. This is the recommended approach for optimizing file discovery in Auto Loader, especially for directories with millions of files.
- ✗
Set cloudFiles.format to 'json'.
Why it's wrong here
Specifying the file format is necessary for parsing but does not affect file discovery performance. It tells Auto Loader how to read the files, not how to find them. The slow ingestion is due to listing overhead, so this option does not solve the problem.
- ✗
Set cloudFiles.maxFilesPerTrigger to a high value.
Why it's wrong here
Increasing maxFilesPerTrigger controls how many files are processed per micro-batch, but it does not improve file discovery speed. It may actually increase per-batch processing time. The bottleneck is listing files, not processing them. Therefore, this option does not address the slow ingestion caused by file discovery.
- ✗
Enable cloudFiles.schemaEvolutionMode to 'rescue'.
Why it's wrong here
This option controls how schema evolution is handled, specifically rescuing data that doesn't match the schema. It has no impact on file discovery performance. The issue is slow listing, not schema handling. Thus, this option is irrelevant to the scenario.
About these practice questions
Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.