Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question
Exhibit
Error: org.apache.spark.SparkException: Job aborted due to stage failure: Total size of serialized results of 5000 tasks (2048 MB) is bigger than spark.driver.maxResultSize (1024 MB)
Refer to the exhibit. What is the best way to resolve this error?
⚠ Common exam trap
Candidates often attempt to increase 'spark.driver.maxResultSize' to fix the error, rather than changing their code pattern to avoid bringing massive data back to the driver node.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Write the results to a data lake instead of collecting them.
This error occurs when the result set returned to the driver from the executors exceeds the limit defined by `spark.driver.maxResultSize`. The best practice is to stop trying to bring huge datasets back to the driver, and instead write the results to a distributed storage system like S3 or ADLS. This prevents the driver from becoming a memory bottleneck and ensures the application remains scalable for large-scale data processing tasks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the spark.driver.maxResultSize setting.
Why it's wrong here
While this avoids the error, it is a band-aid that places massive memory pressure on the driver. The driver node is not designed to act as a storage or aggregation unit for huge datasets, and this approach will eventually lead to driver OOM errors as the data grows over time.
- ✓
Write the results to a data lake instead of collecting them.
Why this is correct
Writing results to persistent, distributed storage (like Parquet files on S3/ADLS) is the architecture-correct way to handle large outputs. This avoids sending all data to the driver, allowing the Spark executors to perform the work in parallel and keeping the driver's memory footprint small and stable throughout the execution.
- ✗
Disable the Spark broadcast join threshold.
Why it's wrong here
Broadcast join settings have no relation to the result size returned to the driver. The error is caused by the final action returning data to the driver, not by join operations during the middle of the job's execution stages. This parameter change will not address the error at all.
- ✗
Increase the number of executor instances.
Why it's wrong here
Adding more executors might speed up the job, but it does not change the amount of data being returned to the driver. If the final result is large, you will still hit the driver result size limit regardless of how many executors were used to compute that final output data.
Visual reference
Quick reference
AWS S3 Storage Class Comparison
| Storage Class | Min Duration | Retrieval | Use Case |
|---|---|---|---|
| S3 Standard | None | Immediate | Frequently accessed data |
| S3 Standard-IA | 30 days | Immediate | Infrequent access, rapid retrieval |
| S3 One Zone-IA | 30 days | Immediate | Non-critical infrequent data |
| S3 Intelligent-Tiering | None | Immediate–hours | Unknown or changing access patterns |
| S3 Glacier Instant | 90 days | Milliseconds | Archive with instant retrieval |
| S3 Glacier Flexible | 90 days | Minutes–hours | Archive, flexible retrieval |
| S3 Glacier Deep Archive | 180 days | Hours | Long-term compliance archive |
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.