Databricks-Spark-Assoc Spark Architecture and Components Practice Question
A Spark job reads a large Parquet dataset, performs a filter, and then a groupBy aggregation. The job's DAG shows two stages: one for the filter and one for the aggregation. The first stage has 200 tasks, and the second stage has 200 tasks. The job is running on a cluster with 10 executors, each with 8 cores. The engineer observes that the second stage takes significantly longer than the first. Which of the following is the most likely cause for the increased duration in the second stage?
⚠ Common exam trap
The trap here is focusing on the number of tasks or partitions as the cause of slowness, rather than recognizing the inherent cost of a shuffle operation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The second stage involves a shuffle, which requires data to be repartitioned across the network, causing additional I/O and serialization overhead.
The groupBy aggregation requires a shuffle, which redistributes data across the cluster based on the grouping key. This involves network transfer, disk I/O, and serialization/deserialization, all of which add significant overhead compared to narrow transformations like filter. Even with the same number of tasks, the shuffle makes the second stage slower.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The second stage involves a shuffle, which requires data to be repartitioned across the network, causing additional I/O and serialization overhead.
Why this is correct
The groupBy aggregation triggers a shuffle, where data is redistributed across executors based on the grouping key. This involves network transfer, disk I/O, and serialization, making the second stage slower than the filter stage, which is narrow and operates on data locally. The shuffle is the primary reason for the increased duration.
- ✗
The second stage is reading from disk, while the first stage reads from memory, causing slower performance.
Why it's wrong here
The first stage reads the initial Parquet data, which may involve disk I/O. The second stage operates on the output of the shuffle, which may also involve disk I/O. The assumption that one reads from memory and the other from disk is not given and is not the main factor; the shuffle itself is the dominant cost.
- ✗
The second stage has a larger number of partitions than the first stage, causing more tasks to be scheduled.
Why it's wrong here
Both stages have 200 tasks, indicating the same number of partitions. The difference in duration is not due to partition count but due to the shuffle operation in the second stage. Increasing partitions would not inherently make the stage slower; it's the data movement that adds time.
- ✗
The second stage has more tasks than the first stage, leading to higher scheduling overhead.
Why it's wrong here
Both stages have 200 tasks, so the number of tasks is identical. The increased duration is not due to task count but due to the nature of the operations within the stage. Scheduling overhead is generally small compared to the cost of a shuffle, so this is not the primary cause.
About these practice questions
Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.