Databricks-Spark-Assoc Spark Architecture and Components Practice Question
A data engineer runs a PySpark job on a Databricks cluster. The job reads a 500 GB Parquet dataset, applies a filter, and writes the result. The engineer notices that during execution, all tasks of a particular stage complete quickly except for a handful that take far longer, and the Spark UI shows these tasks are processing partitions that contain far more records than others. Which Spark architecture concept best explains this behavior, and what is the most appropriate remediation?
⚠ Common exam trap
The trap here is assuming slow tasks always indicate a memory shortage, when uneven partition sizes from a shuffle are the more likely cause in this scenario.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Data skew across partitions during a shuffle; the engineer should apply salting or repartitioning to distribute records more evenly.
Skewed partition sizes after a shuffle produce a few long-running tasks while most finish quickly, which matches the Spark UI pattern described. Redistributing records through salting or repartitioning balances the workload across tasks, addressing the root cause rather than masking symptoms with memory tuning or reduced parallelism.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Insufficient executor memory causing garbage collection pauses only on certain tasks; the engineer should increase spark.executor.memory.
Why it's wrong here
Memory pressure and GC pauses would typically affect many tasks and show up as prolonged GC time in the Spark UI rather than a small number of long-running tasks tied to large partitions. Increasing executor memory does not redistribute records, so the underlying imbalance in partition sizes would remain and the stragglers would persist.
- ✓
Data skew across partitions during a shuffle; the engineer should apply salting or repartitioning to distribute records more evenly.
Why this is correct
Uneven record distribution across partitions after a shuffle causes a few tasks to process disproportionately large partitions, producing stragglers. Salting keys or repartitioning redistributes data so each task receives a comparable workload, directly addressing the imbalance observed in the Spark UI task duration metrics for this stage.
- ✗
The DAG Scheduler is serializing stages incorrectly; the engineer should disable adaptive query execution to force static stage boundaries.
Why it's wrong here
The DAG Scheduler builds stages from shuffle boundaries and does not serialize stages incorrectly in this way. Disabling adaptive query execution removes a feature that can coalesce skewed partitions at runtime, so it would likely make the straggler problem worse instead of resolving the uneven task durations.
- ✗
Too few partitions in the source Parquet files; the engineer should call coalesce(1) before writing the output.
Why it's wrong here
Coalesce(1) reduces parallelism and would concentrate more data into a single task, worsening stragglers rather than fixing them. The described symptom is uneven partition sizes, not an insufficient number of partitions, so merging partitions would not correct the imbalance and could degrade throughput further.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.