Courseiva

Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question

A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?

⚠ Common exam trap

The trap here is assuming that increasing shuffle partitions or enabling AQE will solve small-file problems, when in fact compaction is required.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run OPTIMIZE on the source Delta table to compact small files.

The Spark UI reveals many small tasks and thousands of small files in the source Delta table, which is a small-file problem. OPTIMIZE compacts these small files into larger ones, reducing task overhead and improving read performance. Other options either do not address the root cause or could exacerbate the issue by creating more partitions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable Adaptive Query Execution (AQE) by setting spark.sql.adaptive.enabled to true.

    Why it's wrong here

    AQE is already enabled by default in Databricks and optimizes shuffle partitions and joins dynamically. However, it does not automatically compact small files at the storage layer. While AQE can coalesce partitions, the root cause here is the presence of thousands of small files, which requires file compaction, not just query optimization.

  • ✗

    Increase the number of shuffle partitions by setting spark.sql.shuffle.partitions to a higher value.

    Why it's wrong here

    Increasing shuffle partitions might help with skewed shuffles, but the issue here is small files causing many small tasks, not a lack of parallelism. More shuffle partitions would likely worsen the problem by creating even more, smaller tasks. The Spark UI indicates many small tasks reading small files, so addressing file size is key.

  • ✓

    Run OPTIMIZE on the source Delta table to compact small files.

    Why this is correct

    OPTIMIZE compacts many small files into larger ones, reducing the number of tasks and improving read throughput. In this scenario, the Spark UI shows thousands of small files and many small tasks, which is a classic small-file problem. Compaction directly addresses this by merging files, leading to fewer, larger files that are more efficient to read.

  • ✗

    Repartition the source DataFrame to a smaller number of partitions before writing.

    Why it's wrong here

    Repartitioning the DataFrame affects the write parallelism but does not compact existing small files in the source table. The source table already contains many small files; reading it will still incur overhead. Repartitioning before writing could help future writes, but it does not address the current read performance issue caused by existing small files.

About these practice questions

Courseiva writes every Databricks-DE-Assoc question from scratch — 276 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.