Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
A Spark job reads a large Parquet dataset, performs a groupBy on a high-cardinality column, and writes the result to a Delta table. The job fails with a FetchFailedException on a particular executor. The Spark UI shows that the executor had sufficient memory but the shuffle fetch failed due to a connection reset. Which configuration change is most likely to resolve this issue?
⚠ Common exam trap
The trap here is assuming that memory or query planning changes will fix a shuffle fetch failure, when the error is network-related and requires retry tuning.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set spark.shuffle.io.maxRetries and spark.shuffle.io.retryWait to higher values to handle transient network issues.
A FetchFailedException with connection reset during shuffle fetch is typically caused by transient network issues or shuffle service timeouts. Increasing spark.shuffle.io.maxRetries and spark.shuffle.io.retryWait makes the shuffle fetch more resilient by retrying failed attempts. This is the most direct configuration change to handle temporary network glitches without altering the overall job logic or resource allocation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase spark.executor.memory to prevent the executor from being killed during shuffle.
Why it's wrong here
The scenario states the executor had sufficient memory, so increasing memory would not address the connection reset. The failure is in the shuffle fetch phase, not due to memory pressure. This change would waste resources and not resolve the network-related fetch failure.
- ✗
Increase spark.reducer.maxSizeInFlight to allow larger shuffle blocks to be fetched.
Why it's wrong here
Increasing maxSizeInFlight allows more data to be fetched in parallel, which could increase network pressure and potentially worsen connection resets. The error is a connection reset during shuffle fetch, often caused by the shuffle service being overwhelmed or timing out. Larger fetch sizes may exacerbate the problem rather than solve it.
- ✓
Set spark.shuffle.io.maxRetries and spark.shuffle.io.retryWait to higher values to handle transient network issues.
Why this is correct
FetchFailedException due to connection reset often indicates transient network problems or shuffle service timeouts. Increasing shuffle I/O retries and retry wait allows the reducer to retry fetching blocks after a failure, improving resilience. This directly addresses the connection reset by giving the fetch more attempts and time to succeed.
- ✗
Set spark.sql.adaptive.enabled=false to disable adaptive query execution and avoid shuffle re-computation.
Why it's wrong here
Disabling adaptive query execution would not fix a connection reset during shuffle fetch. AQE may actually help by reducing shuffle size or handling skew. Turning it off could lead to less efficient execution and does not target the network issue. The root cause is transient fetch failure, not query planning.
Visual reference
About these practice questions
This Databricks-Spark-Assoc question is part of Courseiva's 295-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.