Databricks-DE-Assoc Troubleshooting, Monitoring, and Optimization Practice Question
A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?
⚠ Common exam trap
The trap here is assuming that increasing trigger interval or shuffle partitions will solve streaming lag, when the real issue is often small files in the source Delta table.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tune the source Delta table by running OPTIMIZE to compact small files.
For a Structured Streaming query reading from a Delta table, many small files can cause high read overhead per micro-batch, leading to low processing rates and increasing batch durations. Compacting the source table with OPTIMIZE reduces the number of files, improving read efficiency. Other options either do not address the root cause or could exacerbate the issue.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the trigger interval to process larger batches less frequently.
Why it's wrong here
Increasing the trigger interval (e.g., from 10 seconds to 1 minute) would process larger batches, potentially worsening the backlog if the processing rate is already low. It might help with throughput if the bottleneck is per-batch overhead, but the symptom of increasing batch duration suggests the query is falling behind, so larger batches would likely increase latency further.
- ✗
Enable Delta Lake optimized writes by setting spark.databricks.delta.optimizeWrite.enabled to true.
Why it's wrong here
Optimized writes reduce the number of small files written by coalescing data before writing, which improves write performance. However, the issue here is low processing rate and increasing batch duration, which is more likely due to reading inefficiencies or resource contention. Optimized writes address write-side file proliferation, not the processing bottleneck described.
- ✗
Increase the number of shuffle partitions to improve parallelism during processing.
Why it's wrong here
Increasing shuffle partitions can help if the query involves shuffles (e.g., aggregations), but it is not a universal fix. The scenario does not specify a shuffle-heavy operation; the bottleneck could be due to input reading or state management. Without evidence of shuffle spill, this action may not help and could add overhead.
- ✓
Tune the source Delta table by running OPTIMIZE to compact small files.
Why this is correct
A common cause of low processing rate in streaming queries reading from Delta is the accumulation of many small files, which increases the overhead of reading each micro-batch. Running OPTIMIZE compacts these files, reducing the number of files to read per batch and improving throughput. This directly addresses the input side, allowing the query to process data faster and reduce batch duration.
About these practice questions
This Databricks-DE-Assoc question is part of Courseiva's 276-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Assoc exam.