Courseiva
easyMultiple Choice

Optimize Azure Databricks ETL Performance with Delta Lake and Photon

A data engineer needs to process a large dataset stored in Azure Blob Storage using Azure Databricks. The dataset consists of millions of small CSV files. The processing job is slow due to the overhead of reading many small files. Which technique should be used to improve performance?

⚠ Common exam trap

A common mix-up: candidates assume performance issues are always solved by scaling out (Option A) or by switching formats (Option B), but the DP-203 exam specifically tests the understanding that small file overhead is a distinct problem requiring file consolidation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Coalesce the small files into larger files using a Databricks notebook

Coalescing the millions of small CSV files into larger files reduces the metadata overhead and I/O operations when reading from Azure Blob Storage. Databricks can then process fewer, larger files more efficiently, as each task handles a substantial data chunk rather than incurring the cost of opening and closing many small files.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of worker nodes in the cluster

    Why it's wrong here

    Adding worker nodes increases parallel compute, but the bottleneck is per-file open and listing overhead across millions of objects, which more executors cannot remove. Scaling out suits CPU- or memory-bound workloads, not metadata-heavy small-file reads.

  • ✗

    Convert the CSV files to Parquet format

    Why it's wrong here

    Parquet is columnar and compresses well, yet converting millions of tiny files preserves the same object count, so listing and open overhead persists. Format conversion suits reducing storage footprint or enabling column pruning on already-coalesced data, not consolidating small files.

  • ✓

    Coalesce the small files into larger files using a Databricks notebook

    Why this is correct

    Consolidating millions of small CSVs into fewer large files removes per-file listing and open overhead, which is the stated bottleneck. Databricks reads large files far more efficiently, so coalescing directly addresses the small-file problem and speeds up the job.

  • ✗

    Use Delta Lake caching to store the data in memory

    Why it's wrong here

    Delta Lake caching stores already-read data in memory, so it does not reduce the per-file listing and open overhead caused by millions of small CSV files; compaction or batching addresses that. It is tempting because caching accelerates repeated queries over the same dataset, which is a different bottleneck.

Quick reference

Azure Blob Storage Tier Comparison

TierStorage CostRetrieval CostLatencyUse Case
HotHighestLowestImmediateActive data, frequent reads
CoolLowerHigherImmediateData accessed < once / month
ColdLower stillHigherImmediateData accessed < once / quarter
ArchiveLowestHighest + rehydration delayHoursLong-term compliance retention

About these practice questions

Courseiva writes every DP-203 question from scratch — 509 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DP-203 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DP-203 exam.