Courseiva

Databricks-Spark-Assoc Troubleshooting and Tuning DataFrame Apps Practice Question

Exhibit

Error: org.apache.spark.SparkException: Job aborted due to stage failure: Total size of serialized results of 5000 tasks (2048 MB) is bigger than spark.driver.maxResultSize (1024 MB)

Refer to the exhibit. What is the best way to resolve this error?

⚠ Common exam trap

Candidates often attempt to increase 'spark.driver.maxResultSize' to fix the error, rather than changing their code pattern to avoid bringing massive data back to the driver node.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Write the results to a data lake instead of collecting them.

This error occurs when the result set returned to the driver from the executors exceeds the limit defined by `spark.driver.maxResultSize`. The best practice is to stop trying to bring huge datasets back to the driver, and instead write the results to a distributed storage system like S3 or ADLS. This prevents the driver from becoming a memory bottleneck and ensures the application remains scalable for large-scale data processing tasks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the spark.driver.maxResultSize setting.

    Why it's wrong here

    While this avoids the error, it is a band-aid that places massive memory pressure on the driver. The driver node is not designed to act as a storage or aggregation unit for huge datasets, and this approach will eventually lead to driver OOM errors as the data grows over time.

  • ✓

    Write the results to a data lake instead of collecting them.

    Why this is correct

    Writing results to persistent, distributed storage (like Parquet files on S3/ADLS) is the architecture-correct way to handle large outputs. This avoids sending all data to the driver, allowing the Spark executors to perform the work in parallel and keeping the driver's memory footprint small and stable throughout the execution.

  • ✗

    Disable the Spark broadcast join threshold.

    Why it's wrong here

    Broadcast join settings have no relation to the result size returned to the driver. The error is caused by the final action returning data to the driver, not by join operations during the middle of the job's execution stages. This parameter change will not address the error at all.

  • ✗

    Increase the number of executor instances.

    Why it's wrong here

    Adding more executors might speed up the job, but it does not change the amount of data being returned to the driver. If the final result is large, you will still hit the driver result size limit regardless of how many executors were used to compute that final output data.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

Quick reference

AWS S3 Storage Class Comparison

Storage ClassMin DurationRetrievalUse Case
S3 StandardNoneImmediateFrequently accessed data
S3 Standard-IA30 daysImmediateInfrequent access, rapid retrieval
S3 One Zone-IA30 daysImmediateNon-critical infrequent data
S3 Intelligent-TieringNoneImmediate–hoursUnknown or changing access patterns
S3 Glacier Instant90 daysMillisecondsArchive with instant retrieval
S3 Glacier Flexible90 daysMinutes–hoursArchive, flexible retrieval
S3 Glacier Deep Archive180 daysHoursLong-term compliance archive

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.