Databricks-Spark-Assoc Spark Architecture and Components Practice Question
A developer is using Spark on Databricks and notices that a particular job has many stages due to shuffle operations. They want to understand the role of the shuffle in the Spark execution model. Which two statements accurately describe the behavior of a shuffle operation in Spark? (Choose two.)
⚠ Common exam trap
The trap here is assuming that a shuffle guarantees a one-to-one mapping between map tasks and reduce partitions, or that actions directly cause shuffles.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A shuffle writes intermediate data to disk on the map side and reads it over the network on the reduce side.
A shuffle operation writes intermediate data to disk on the map side and transfers it over the network to reduce tasks, creating a stage boundary. These two characteristics are central to understanding why shuffles are expensive and how Spark's DAG is structured. The number of reduce partitions is configurable and not fixed to one per reducer, and shuffles are triggered by transformations, not actions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A shuffle eliminates the need for the driver to coordinate task scheduling across stages.
Why it's wrong here
This is incorrect. The driver, specifically the DAG Scheduler, still coordinates stage execution and task scheduling. In fact, shuffles introduce additional coordination because the DAG Scheduler must wait for map stages to complete before submitting reduce stages. The shuffle service on executors handles data transfer, but the driver orchestrates the overall job flow, including stage retries if shuffle outputs are lost.
- ✓
A shuffle writes intermediate data to disk on the map side and reads it over the network on the reduce side.
Why this is correct
This is correct. During a shuffle, map tasks write shuffle files to local disk (or memory if configured) and reduce tasks fetch these files over the network. This disk I/O and network transfer make shuffles expensive. Spark's shuffle manager (e.g., SortShuffleManager) handles this process, and the data is partitioned by key before being written, ensuring that all records for a given key end up on the same reducer.
- ✓
A shuffle creates a stage boundary, splitting the job into a map stage and a reduce stage.
Why this is correct
This is correct. A shuffle dependency is a wide dependency that forces a new stage. The upstream stage (map stage) produces shuffle data, and the downstream stage (reduce stage) consumes it. This boundary is fundamental to Spark's DAG scheduling: stages are separated by shuffles, and the DAG Scheduler submits them in topological order, ensuring that map outputs are available before reduce tasks start.
- ✗
A shuffle always results in exactly one partition per reducer task, regardless of the number of map tasks.
Why it's wrong here
This is incorrect. The number of reduce partitions is determined by the `spark.sql.shuffle.partitions` (for SQL/DataFrame) or the `numPartitions` parameter of the transformation (e.g., `groupByKey(numPartitions)`). It does not automatically match the number of map tasks. The number of map tasks (and thus map output files) can be different, and each reducer fetches only its relevant partitions from all map outputs.
- ✗
A shuffle is only triggered by actions, not by transformations.
Why it's wrong here
This is incorrect. Shuffles are triggered by certain transformations that require data redistribution, such as `groupByKey`, `reduceByKey`, `join`, and `repartition`. Actions like `count` or `collect` trigger job execution but do not directly cause shuffles; they simply force the evaluation of the DAG, which may include shuffle transformations. Thus, shuffles are a property of transformations, not actions.
About these practice questions
One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.