Courseiva

Databricks-Spark-Assoc Spark Architecture and Components Practice Question

A developer is tuning a Databricks job and wants to know how many tasks will be created for the final stage of a job that reads a Parquet file with 200 partitions, applies a filter, and then calls coalesce(10) before writing the result. Assuming no other repartitioning or shuffles occur, how many tasks will the final write stage contain?

⚠ Common exam trap

Candidates often confuse coalesce with repartition, or assuming that filter triggers a shuffle, when in fact coalesce only merges partitions without a full shuffle and filter is narrow.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

10 tasks, because coalesce(10) reduces the partition count to 10

coalesce(10) reduces the RDD from 200 partitions to 10 using narrow dependencies, so no shuffle occurs. Since the number of tasks in a stage equals the number of partitions in the RDD being processed, the final write stage launches 10 tasks. Filtering is also narrow and does not alter the partition count, so the coalesce result directly determines the task count.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    1 task, because coalesce always merges everything into a single partition

    Why it's wrong here

    coalesce only reduces to the number specified as its argument; it does not default to one partition unless the argument is 1. Passing 10 explicitly requests 10 partitions, so the final stage would have 10 tasks, not a single task. The misunderstanding that coalesce always produces one partition is a common misconception.

  • ✗

    200 tasks, because the filter forces a shuffle before coalesce

    Why it's wrong here

    A filter transformation is narrow; it does not trigger a shuffle because each output partition depends only on its corresponding input partition. Therefore, no shuffle occurs before coalesce, and the partition count after coalesce(10) is 10. The claim that filter forces a shuffle misrepresents Spark's narrow versus wide dependency classification.

  • ✗

    200 tasks, because coalesce does not change the number of partitions

    Why it's wrong here

    coalesce does reduce the number of partitions by merging existing ones without a full shuffle, so claiming it leaves 200 partitions is incorrect. The narrow dependency of coalesce allows Spark to combine partitions, and the resulting RDD has the reduced count, which directly determines the number of tasks in the following stage.

  • ✓

    10 tasks, because coalesce(10) reduces the partition count to 10

    Why this is correct

    coalesce(10) collapses the 200 input partitions into 10 output partitions using narrow dependencies, avoiding a shuffle. Each partition corresponds to one task in the stage that writes the data, so the final stage launches exactly 10 tasks. This matches the intent of coalesce: reducing partition count efficiently without redistributing data across the cluster.

About these practice questions

Courseiva writes every Databricks-Spark-Assoc question from scratch — 295 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.