Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
Exhibit
24/05/10 10:00:00 WARN DAGScheduler: Broadcasting large task binary with size 12.5 MiB
Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?
⚠ Common exam trap
Candidates often assume the solution is to increase the executor memory. However, the error is caused by serializing a large object in a closure, which remains an issue regardless of total heap size.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A large variable is captured in a closure; use a Broadcast Variable.
The warning indicates that a large object, likely a high-dimensional collection or a large variable, is being captured in a closure and broadcast to all executors. This increases network pressure and memory usage. The mitigation is to use a broadcast variable or remove the reference to the large object from the closure to prevent serialization overhead and potential performance degradation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The partition size is too large; increase spark.sql.files.maxPartitionBytes.
Why it's wrong here
This setting controls how much data a partition can hold when reading files. It has no relation to the task binary size, which is caused by objects captured within code closures. Increasing this value would only change the number of files processed, not the serialization behavior.
- ✓
A large variable is captured in a closure; use a Broadcast Variable.
Why this is correct
When a large object is referenced inside a transformation, Spark tries to serialize it with the task. A broadcast variable provides a mechanism to distribute the object efficiently once per node rather than once per task, preventing the warning and reducing the serialization burden on the driver.
- ✗
The executors have insufficient memory; increase spark.executor.memory.
Why it's wrong here
While more memory might stop an OOM error, it does not fix the underlying inefficiency of sending large objects inside task binaries. The warning specifically points to a coding pattern that captures excessive data in closures, which should be refactored regardless of how much memory is available.
- ✗
The cluster is out of network bandwidth; enable compression.
Why it's wrong here
While compression can reduce network usage, it does not address the fundamental issue of capturing a large object within a closure. The warning is a symptom of poor code architecture where the developer is inadvertently causing large objects to be serialized and sent to every task.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.