Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?
⚠ Common exam trap
The trap here is treating a slowly growing streaming batch time as a parallelism problem and adding shuffle partitions instead of investigating state accumulation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check for stateful aggregation or deduplication without a watermark and add a watermark with a bounded state cleanup so state does not grow unbounded.
A steady increase in batch duration over hours with a compacted source and adequate CPU strongly indicates unbounded state growth in a stateful streaming operation. Without a watermark, Spark retains all keys indefinitely, so each micro-batch reads and writes a larger state store. Adding a watermark with a bounded delay lets Spark evict old state, keeping per-batch work roughly constant and restoring stable batch times.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase spark.sql.shuffle.partitions to a very high value so each micro-batch uses more tasks and finishes faster.
Why it's wrong here
Raising shuffle partitions creates many tiny tasks and more scheduling overhead per micro-batch, which typically increases, not decreases, batch duration. It also multiplies the number of small output files in the sink, worsening the small-file problem. The bottleneck described is not a shortage of parallelism, since CPU is sufficient and the source is compacted.
- ✗
Reduce the number of shuffle partitions and disable adaptive query execution so the streaming plan stays deterministic across micro-batches.
Why it's wrong here
Disabling adaptive query execution removes runtime optimizations such as coalescing shuffle partitions and skew join handling, which are helpful for streaming micro-batches. Reducing shuffle partitions alone does not address growing state or metadata. This option is a step backward and does not explain the gradual slowdown.
- ✗
Enable Delta Lake optimized writes and tune the trigger interval so the job processes larger, less frequent micro-batches.
Why it's wrong here
Optimized writes reduce file count and are beneficial, but increasing the trigger interval does not fix a growing batch duration caused by accumulating state or metadata. Larger micro-batches would process more data per trigger and could make each batch even slower. The root cause of the growth is state or log accumulation, not the trigger cadence alone.
- ✓
Check for stateful aggregation or deduplication without a watermark and add a watermark with a bounded state cleanup so state does not grow unbounded.
Why this is correct
Unbounded state from a stateful operation without a watermark causes each micro-batch to scan an ever-growing state store, which steadily increases batch duration and can make the job fall behind. Adding a watermark lets Spark drop state older than the watermark, keeping state bounded and batch times stable. This directly addresses a gradual, hours-long degradation pattern.
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.