Courseiva

Databricks-Spark-Assoc Spark Architecture and Components Practice Question

A data engineer is tuning a Spark Structured Streaming job on Databricks that reads from a Kafka topic with 12 partitions. The job uses a static allocation of executors, each with 4 cores. The engineer notices that only 4 tasks are running concurrently, even though there are 12 Kafka partitions and 3 executors are available. Which Spark configuration is most likely causing this limitation?

⚠ Common exam trap

The trap here is assuming that the number of input partitions directly dictates concurrency, ignoring the executor allocation configuration that caps the available cores.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

spark.executor.instances is set to 1, so only one executor with 4 cores is available.

The concurrency of tasks in Spark is determined by the total number of cores available across all executors. If only one executor with 4 cores is allocated, only 4 tasks can run in parallel, regardless of the number of input partitions. The configuration spark.executor.instances controls how many executors are requested, and setting it to 1 restricts the job to a single executor's worth of cores.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    spark.executor.cores is set to 4, limiting each executor to 4 concurrent tasks.

    Why it's wrong here

    With 3 executors each having 4 cores, the total available cores would be 12, which should allow 12 concurrent tasks. The observed concurrency of 4 indicates that only one executor is being used, not that each executor's core count is limiting. This setting alone does not explain why only 4 tasks run when 12 partitions exist.

  • ✗

    spark.default.parallelism is set to 4, overriding the number of partitions from Kafka.

    Why it's wrong here

    spark.default.parallelism influences the number of partitions for RDDs created by transformations like join or reduceByKey, but for Structured Streaming sources like Kafka, the number of partitions is determined by the source itself (e.g., Kafka partitions). It does not override Kafka partition count, so it would not cap concurrency at 4 in this scenario.

  • ✗

    spark.sql.shuffle.partitions is set to 4, limiting the number of tasks for shuffle operations.

    Why it's wrong here

    spark.sql.shuffle.partitions controls the number of partitions created after a shuffle (e.g., for aggregations or joins). While it can limit parallelism for shuffle stages, the initial reading from Kafka is not a shuffle operation; it uses the Kafka partitions directly. Thus, it would not restrict the initial concurrency to 4 tasks.

  • ✓

    spark.executor.instances is set to 1, so only one executor with 4 cores is available.

    Why this is correct

    If spark.executor.instances is set to 1, only a single executor with 4 cores is launched, allowing at most 4 concurrent tasks. Despite having 3 executors' worth of resources potentially available, the static allocation configuration limits the job to one executor. This directly explains the observed concurrency of 4 tasks, matching the executor's core count.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.