Databricks-DE-Assoc Data Transformation and Modeling Practice Question
A data engineer is building a Databricks job that processes millions of small JSON files landed in cloud storage each hour. The job currently spends most of its runtime on file listing and task scheduling overhead. The engineer wants to improve throughput without changing the downstream table schema. Which change should be made to the ingestion step?
⚠ Common exam trap
The trap here is assuming that a target-side optimization such as Auto Compaction solves a source-side file listing bottleneck.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Auto Loader with cloudFiles and enable file notification mode to ingest new files incrementally.
The bottleneck is file discovery and task scheduling for many small files, not the target table layout or columnar format. Auto Loader with cloudFiles incrementally discovers new files, and file notification mode avoids full directory listings by using cloud storage event notifications, which scales to high file counts. This preserves the existing schema while improving ingestion throughput.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the source JSON files to Parquet with a separate Spark batch job before each hourly run.
Why it's wrong here
A separate conversion job still has to list and read all the small JSON files, so the overhead is merely moved earlier. It also adds latency and another failure point in the pipeline. While Parquet is more efficient for downstream reads, the scenario requires improving the ingestion step itself, which this approach does not do.
- ✗
Enable Auto Compaction on the target Delta table so small files are merged after each write.
Why it's wrong here
Auto Compaction reduces small files in the target Delta table after writes, which helps downstream reads, but it does not address the overhead of listing and scheduling millions of tiny source JSON files. The ingestion step still pays the file discovery and task launch cost, so this setting does not improve the throughput of the initial read.
- ✓
Use Auto Loader with cloudFiles and enable file notification mode to ingest new files incrementally.
Why this is correct
Auto Loader with cloudFiles tracks newly arrived files using a scalable discovery mechanism, and file notification mode uses cloud queue notifications instead of repeated directory listings. This directly removes the listing and scheduling overhead described, while schema inference and evolution keep the downstream Delta table schema stable.
- ✗
Repartition the DataFrame to one partition per input file before writing to the target table.
Why it's wrong here
Creating one partition per input file multiplies task count and scheduling pressure, which worsens the exact bottleneck. It also creates many tiny output files in the target table. Repartitioning changes parallelism distribution but does nothing to reduce source file listing or task launch cost, so throughput will not improve.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.