Courseiva

Databricks-Spark-Assoc Spark Architecture and Components Practice Question

Exhibit

Task 12: 500MB (Local Disk)
Task 15: 1.2GB (Shuffle Read)
Task 22: 400MB (Shuffle Write)

Refer to the exhibit. Which performance indicator suggests that Task 15 is likely causing a performance bottleneck during the execution of a join operation?

⚠ Common exam trap

Candidates often misread task duration or spill metrics as the primary indicator of skew, ignoring the distinct volume imbalance shown by massive shuffle read sizes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Task 15 having 1.2GB Shuffle Read.

The 'Shuffle Read' metric indicates the volume of data transferred over the network to that specific task. A high shuffle read compared to other tasks often points to data skew, where one partition receives significantly more data than others. This is a critical insight for developers because skew can lead to uneven executor loads, where one executor works much longer than others, effectively slowing down the entire stage of the Spark job.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Task 12 having 500MB on Local Disk.

    Why it's wrong here

    Disk usage is normal for tasks that spill data or manage intermediate results. Since 500MB is relatively small compared to the 1.2GB shuffle read seen in other tasks, this does not represent a significant bottleneck or an unusual distribution of work that would cause a major performance delay.

  • ✓

    Task 15 having 1.2GB Shuffle Read.

    Why this is correct

    A high shuffle read value suggests that the task is processing a disproportionately large amount of data compared to its peers. This is a classic sign of data skew, where a specific key is overrepresented, causing the executor assigned to that partition to take much longer to finish.

  • ✗

    Task 22 having 400MB Shuffle Write.

    Why it's wrong here

    Shuffle write represents data being output to the next stage. A value of 400MB is standard and expected for a partition's output. It does not indicate a bottleneck because it represents the completion of work rather than the reception of an overwhelming amount of data from other nodes.

  • ✗

    The total volume across all tasks is too low.

    Why it's wrong here

    The total volume is irrelevant to the existence of a bottleneck within a single task. A bottleneck occurs when work is unbalanced, not when the overall dataset is too small. Even with small datasets, a skewed partition will cause the specific task processing it to lag behind others.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.