Courseiva

Databricks-Spark-Assoc · domain

Spark Architecture and Components

This domain covers how a Spark application is structured and executed on Databricks: the driver, executors, cluster manager, and the difference between transformations and actions. Questions test your understanding of what triggers job execution, how deploy modes and resource allocation work, and which component owns scheduling, state, and task execution.

69 questions17 easy35 medium17 hard

Focused practice

Practice Spark Architecture and Components questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Spark Architecture and Components

Be able to trace what happens from a DataFrame action to jobs, stages, and tasks, and identify which component does what. The key point: the driver builds the plan and schedules work, while executors run tasks and hold cached data.

What happens on the driver when an action is called on a DataFrame

Driver versus executor responsibilities for scheduling, state, and task execution

Client versus cluster deploy mode behavior when submitting with spark-submit

Executor role in caching data, running tasks, and reporting results to the driver

Watch out for

Common Spark Architecture and Components exam traps

  • ▸Assuming transformations like groupBy execute immediately; they are lazy and only build a plan until an action runs.
  • ▸Confusing the driver with the cluster manager: the driver coordinates the application, while the cluster manager allocates resources.
  • ▸Believing executors persist after the application ends or that the driver runs on a worker node in client deploy mode.

Question index

All Spark Architecture and Components questions (69)

Click any question to see the full explanation, or start a practice session above.

1

In Spark's cluster architecture, what happens to the tasks if the Driver node crashes during the execution of a job?

Easy
2

Refer to the exhibit. Which configuration issue is most likely causing this warning in a Spark application?

Hard
3

A Spark application running on Databricks uses broadcast joins for a small dimension table. A developer notices that the broadcast variable is not being sent to executors as expected, causing a shuffle instead. Which configuration property directly controls the maximum size of a table that Spark will automatically broadcast?

Hard
4

Which term describes the unit of work that is dispatched by the Driver to a specific Executor?

Easy
5

A Spark application on Databricks uses a broadcast variable to distribute a large lookup table to all executors. The developer notices that the broadcast variable is not being used efficiently, as executors are still fetching the data multiple times. Which component is responsible for ensuring that the broadcast data is distributed only once per executor and cached there?

Medium
6

When a Spark application is running in Databricks, what determines the number of tasks that can run in parallel?

Medium
7

In the context of Databricks, what occurs when a Spark stage is described as 'Shuffle-heavy'?

Hard
8

A data engineer is tuning a Spark Structured Streaming job on Databricks that reads from a Kafka topic with 12 partitions. The job uses a static allocation of executors, each with 4 cores. The engineer notices that only 4 tasks are running concurrently, even though there are 12 Kafka partitions and 3 executors are available. Which Spark configuration is most likely causing this limitation?

Medium
9

In the context of the Spark Driver, which component is specifically responsible for tracking the location of cached data blocks across the executors?

Medium
10

Which of the following best describes the purpose of a Spark Session in a Databricks environment?

Easy
11

A data engineer submits a PySpark job that performs a wide transformation via a join operation across two large datasets. During execution, several tasks in the shuffle stage fail repeatedly due to transient network timeouts between worker nodes. How does Apache Spark's architecture handle these failed tasks?

Medium
12

A Spark application running on a Databricks cluster uses a broadcast variable to distribute a small lookup table to all executors. During execution, the driver serializes the broadcast variable and sends it to each executor. Which component is responsible for storing the broadcast data on the executor side and making it available to tasks?

Hard
13

Which THREE components are part of the Spark execution environment that resides on the Driver node?

Medium
14

A Spark job reads a large Parquet file, performs a groupBy operation, and then writes the result. During execution, the job fails with an OutOfMemoryError on the Driver. Which component is most likely responsible for the memory issue?

Hard
15

A developer submits a Spark application to a Databricks cluster. The application creates a SparkSession, reads a CSV file, and calls count() on the resulting DataFrame. Which component is responsible for translating this logical operation into a physical execution plan and coordinating its execution across the cluster?

Easy
16

A developer runs a Spark application on a Databricks cluster in Standard access mode. The application reads a Parquet file, applies a filter, and calls `df.cache()` before an action. During execution, the driver logs show that a stage is retried because a task failed with an executor lost error. Which component is responsible for rescheduling the failed task on another executor within the same application?

Medium
17

A developer is writing a Spark application that will run on a Databricks cluster. They need to ensure that the driver program can communicate with the executors and that tasks are distributed correctly. Which component is responsible for coordinating the execution of tasks across the executors?

Easy
18

Refer to the exhibit. Which performance indicator suggests that Task 15 is likely causing a performance bottleneck during the execution of a join operation?

Hard
19

A Spark job reads a large CSV file, performs a groupBy aggregation, and then writes the result. The Spark UI shows that the job has multiple stages, and one stage has a large number of tasks. Which factor primarily determines the number of tasks in the stage that performs the aggregation?

Medium
20

A Spark application is submitted to a Databricks cluster. The application uses a broadcast variable to distribute a small lookup table to all Executors. Which component is responsible for broadcasting this variable?

Medium
21

Which of the following correctly describes the relationship between a Spark Job and a Spark Stage?

Medium
22

Which component in the Spark architecture is responsible for scheduling tasks and managing the execution of jobs on the cluster?

Medium
23

What happens when a Spark job triggers a 'shuffle' operation during execution?

Medium
24

A Spark application is running on a Databricks cluster with 3 worker nodes, each having 4 cores. The application uses the default configuration. How many tasks can run concurrently across the cluster?

Easy
25

A developer is using Spark on Databricks and wants to monitor the progress of a job. They need to understand how the driver coordinates with executors. Which component is responsible for scheduling tasks onto executors and tracking their status?

Easy
26

What is the primary benefit of the Catalyst Optimizer in the Spark SQL architecture?

Medium
27

What is the primary function of the 'Shuffle Service' in a Spark cluster when using dynamic allocation?

Hard
28

What is the purpose of the 'Broadcast Variable' in the Spark architecture?

Medium
29

In a Spark application running on a Databricks cluster, the driver program creates a SparkSession and defines a series of transformations. When an action is triggered, the driver requests resources from the cluster manager. Which component is responsible for negotiating and acquiring these resources on behalf of the Spark application?

Easy
30

A Spark job is running on Databricks and experiences a stage where tasks are taking much longer than expected. The Spark UI shows that some tasks have significantly higher shuffle read sizes than others, and the stage is skewed. Which Spark feature can automatically mitigate this skew by splitting large partitions into smaller ones?

Hard
31

Refer to the exhibit. Based on the error log, what is the most likely cause of the job failure?

Hard
32

A developer is using Spark on Databricks and notices that a particular job has many stages due to shuffle operations. They want to understand the role of the shuffle in the Spark execution model. Which two statements accurately describe the behavior of a shuffle operation in Spark? (Choose two.)

Medium
33

A data engineer runs a PySpark job on a Databricks cluster. The job reads a 500 GB Parquet dataset, applies a filter, and writes the result. The engineer notices that during execution, all tasks of a particular stage complete quickly except for a handful that take far longer, and the Spark UI shows these tasks are processing partitions that contain far more records than others. Which Spark architecture concept best explains this behavior, and what is the most appropriate remediation?

Medium
34

A developer is debugging a Spark job and observes that a particular stage has 200 tasks, but only 10 executors with 2 cores each are available. What will happen to the remaining tasks in that stage?

Medium
35

A developer is tuning a Databricks job and wants to know how many tasks will be created for the final stage of a job that reads a Parquet file with 200 partitions, applies a filter, and then calls coalesce(10) before writing the result. Assuming no other repartitioning or shuffles occur, how many tasks will the final write stage contain?

Medium
36

Which TWO factors influence the effective parallelism of a Spark application?

Medium
37

What is the primary function of the Spark DAG Scheduler?

Medium
38

A data engineer is configuring a Spark application on Databricks. They set `spark.executor.instances` to 4, `spark.executor.cores` to 5, and `spark.executor.memory` to 16g. The cluster has 5 worker nodes, each with 16 cores and 64 GB RAM. What is the maximum number of tasks that can run concurrently across all executors?

Easy
39

A data engineer observes that a Spark Structured Streaming job on Databricks processes micro-batches with steadily increasing latency over several hours. The Spark UI shows that the number of active tasks per batch stays constant, but each task processes a growing amount of state. Which architectural behavior explains this pattern?

Hard
40

Which TWO factors contribute to the 'Data Locality' optimization in Spark?

Medium
41

Which THREE components are involved in the process of executing a Shuffle operation?

Medium
42

A developer is using Spark on Databricks and wants to understand how the Driver and Executors communicate during a job. Which two statements accurately describe this interaction? (Choose two.)

Medium
43

Which TWO of the following statements correctly describe the role of the Spark Executor in a Databricks environment?

Medium
44

Which Spark configuration property determines the maximum amount of memory the Spark Driver can request for itself when running on a Kubernetes cluster?

Medium
45

A Spark job reads a large Parquet dataset, performs a filter, and then a groupBy aggregation. The job's DAG shows two stages: one for the filter and one for the aggregation. The first stage has 200 tasks, and the second stage has 200 tasks. The job is running on a cluster with 10 executors, each with 8 cores. The engineer observes that the second stage takes significantly longer than the first. Which of the following is the most likely cause for the increased duration in the second stage?

Hard
46

A Databricks engineer is diagnosing why a Spark job's shuffle phase writes a very large amount of data to disk. The engineer wants to reduce shuffle overhead by changing how the job is structured and configured. Which TWO actions are most likely to reduce the volume of shuffle data written? (Choose two.)

Hard
47

In the Databricks Spark environment, what is the role of the 'Shuffle Service'?

Medium
48

Which component manages the lifecycle and allocation of executors in a Databricks cluster?

Easy
49

What happens when an action is called on a Spark DataFrame?

Easy
50

A Spark application is running in cluster mode on Databricks. The driver program is running on a worker node, and the application has been running for several hours. Suddenly, the driver node experiences a hardware failure and crashes. What happens to the running tasks and the application?

Hard
51

Which component in the Spark architecture is responsible for maintaining the state of the Spark application and coordinating the execution of tasks across the cluster?

Medium
52

Which property of RDDs (Resilient Distributed Datasets) is primarily responsible for Spark's fault tolerance during cluster execution?

Hard
53

A Spark application is running on a cluster with 5 executors. The driver program creates a broadcast variable that is used in a transformation. Which two components are directly involved in distributing and using the broadcast variable? (Choose two.)

Medium
54

A data engineer submits a Spark application using spark-submit in client deploy mode from an edge node. The application reads a large Parquet dataset, performs a groupBy aggregation, and writes the result to a Delta table. The engineer notices that the Driver process runs on the edge node and remains alive throughout the application's lifetime. Which statement best describes the role of the Driver in this scenario?

Medium
55

Which TWO of the following statements accurately describe the role of the Spark Executor in a cluster deployment?

Hard
56

What is the primary role of the 'Cluster Manager' in Spark?

Medium
57

A developer submits a Spark application to a Databricks cluster using spark-submit with deploy mode set to cluster. During execution, one of the worker nodes hosting a task fails and is lost by the cluster manager. Which Spark component is responsible for rescheduling the failed task on another available executor?

Easy
58

Which TWO of the following statements accurately describe the relationship between Spark Executors and memory management within a Databricks cluster?

Hard
59

What is the primary role of the 'Executor' process in the Spark distributed architecture?

Easy
60

Which component in the Spark cluster architecture is responsible for communicating directly with the Cluster Manager (e.g., YARN, Mesos, K8s) to request and release resources?

Easy
61

Which process is responsible for tracking the location of data blocks cached in the executors?

Medium
62

Which Spark component is responsible for maintaining the Directed Acyclic Graph (DAG) of stages and tasks?

Easy
63

In a Databricks Spark cluster, which component is primarily responsible for scheduling tasks and managing the distribution of computation across the worker nodes?

Medium
64

What is the consequence of having 'wide dependencies' in a Spark job regarding the Spark Architecture?

Hard
65

A data engineer runs a Spark job on a Databricks cluster using the default FIFO scheduler. They notice that a long-running job is holding all cluster resources, and short ad-hoc queries submitted later are stuck waiting. The engineer wants to allow concurrent scheduling of multiple jobs within the same Spark application so that short jobs can run while the long job is still executing. Which Spark configuration should be set to enable this behavior?

Medium
66

In the context of the Spark Driver, what is the 'DAG' and why is it important?

Easy
67

What is the function of the 'Executor' within the Spark execution model?

Medium
68

Refer to the exhibit. Which of the following is the most likely cause for this 'shuffle fetch failure' in a Databricks cluster?

Medium
69

Which of the following describes the 'Driver' process in a Spark application?

Easy

Frequently asked questions

What does the Spark Architecture and Components domain cover on the Databricks-Spark-Assoc exam?
Be able to trace what happens from a DataFrame action to jobs, stages, and tasks, and identify which component does what. The key point: the driver builds the plan and schedules work, while executors run tasks and hold cached data.
How many questions are in this domain?
This page lists all 69 Spark Architecture and Components questions in the Databricks-Spark-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Spark Architecture and Components questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-spark-developer-associate DATABRICKS-SPARK-DEVELOPER-ASSOCIATE spark architecture components Practice Questions