Databricks-Spark-Assoc · domain
Spark Architecture and Components
This domain covers how a Spark application is structured and executed on Databricks: the driver, executors, cluster manager, and the difference between transformations and actions. Questions test your understanding of what triggers job execution, how deploy modes and resource allocation work, and which component owns scheduling, state, and task execution.
Focused practice
Practice Spark Architecture and Components questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Spark Architecture and Components
Be able to trace what happens from a DataFrame action to jobs, stages, and tasks, and identify which component does what. The key point: the driver builds the plan and schedules work, while executors run tasks and hold cached data.
What happens on the driver when an action is called on a DataFrame
Driver versus executor responsibilities for scheduling, state, and task execution
Client versus cluster deploy mode behavior when submitting with spark-submit
Executor role in caching data, running tasks, and reporting results to the driver
Watch out for
Common Spark Architecture and Components exam traps
- ▸Assuming transformations like groupBy execute immediately; they are lazy and only build a plan until an action runs.
- ▸Confusing the driver with the cluster manager: the driver coordinates the application, while the cluster manager allocates resources.
- ▸Believing executors persist after the application ends or that the driver runs on a worker node in client deploy mode.
Question index
All Spark Architecture and Components questions (69)
Click any question to see the full explanation, or start a practice session above.
In Spark's cluster architecture, what happens to the tasks if the Driver node crashes during the execution of a job?
Easy2Refer to the exhibit. Which configuration issue is most likely causing this warning in a Spark application?
Hard3A Spark application running on Databricks uses broadcast joins for a small dimension table. A developer notices that the broadcast variable is not being sent to executors as expected, causing a shuffle instead. Which configuration property directly controls the maximum size of a table that Spark will automatically broadcast?
Hard4Which term describes the unit of work that is dispatched by the Driver to a specific Executor?
Easy5A Spark application on Databricks uses a broadcast variable to distribute a large lookup table to all executors. The developer notices that the broadcast variable is not being used efficiently, as executors are still fetching the data multiple times. Which component is responsible for ensuring that the broadcast data is distributed only once per executor and cached there?
Medium6When a Spark application is running in Databricks, what determines the number of tasks that can run in parallel?
Medium7In the context of Databricks, what occurs when a Spark stage is described as 'Shuffle-heavy'?
Hard8A data engineer is tuning a Spark Structured Streaming job on Databricks that reads from a Kafka topic with 12 partitions. The job uses a static allocation of executors, each with 4 cores. The engineer notices that only 4 tasks are running concurrently, even though there are 12 Kafka partitions and 3 executors are available. Which Spark configuration is most likely causing this limitation?
Medium9In the context of the Spark Driver, which component is specifically responsible for tracking the location of cached data blocks across the executors?
Medium10Which of the following best describes the purpose of a Spark Session in a Databricks environment?
Easy11A data engineer submits a PySpark job that performs a wide transformation via a join operation across two large datasets. During execution, several tasks in the shuffle stage fail repeatedly due to transient network timeouts between worker nodes. How does Apache Spark's architecture handle these failed tasks?
Medium12A Spark application running on a Databricks cluster uses a broadcast variable to distribute a small lookup table to all executors. During execution, the driver serializes the broadcast variable and sends it to each executor. Which component is responsible for storing the broadcast data on the executor side and making it available to tasks?
Hard13Which THREE components are part of the Spark execution environment that resides on the Driver node?
Medium14A Spark job reads a large Parquet file, performs a groupBy operation, and then writes the result. During execution, the job fails with an OutOfMemoryError on the Driver. Which component is most likely responsible for the memory issue?
Hard15A developer submits a Spark application to a Databricks cluster. The application creates a SparkSession, reads a CSV file, and calls count() on the resulting DataFrame. Which component is responsible for translating this logical operation into a physical execution plan and coordinating its execution across the cluster?
Easy16A developer runs a Spark application on a Databricks cluster in Standard access mode. The application reads a Parquet file, applies a filter, and calls `df.cache()` before an action. During execution, the driver logs show that a stage is retried because a task failed with an executor lost error. Which component is responsible for rescheduling the failed task on another executor within the same application?
Medium17A developer is writing a Spark application that will run on a Databricks cluster. They need to ensure that the driver program can communicate with the executors and that tasks are distributed correctly. Which component is responsible for coordinating the execution of tasks across the executors?
Easy18Refer to the exhibit. Which performance indicator suggests that Task 15 is likely causing a performance bottleneck during the execution of a join operation?
Hard19A Spark job reads a large CSV file, performs a groupBy aggregation, and then writes the result. The Spark UI shows that the job has multiple stages, and one stage has a large number of tasks. Which factor primarily determines the number of tasks in the stage that performs the aggregation?
Medium20A Spark application is submitted to a Databricks cluster. The application uses a broadcast variable to distribute a small lookup table to all Executors. Which component is responsible for broadcasting this variable?
Medium21Which of the following correctly describes the relationship between a Spark Job and a Spark Stage?
Medium22Which component in the Spark architecture is responsible for scheduling tasks and managing the execution of jobs on the cluster?
Medium23What happens when a Spark job triggers a 'shuffle' operation during execution?
Medium24A Spark application is running on a Databricks cluster with 3 worker nodes, each having 4 cores. The application uses the default configuration. How many tasks can run concurrently across the cluster?
Easy25A developer is using Spark on Databricks and wants to monitor the progress of a job. They need to understand how the driver coordinates with executors. Which component is responsible for scheduling tasks onto executors and tracking their status?
Easy26What is the primary benefit of the Catalyst Optimizer in the Spark SQL architecture?
Medium27What is the primary function of the 'Shuffle Service' in a Spark cluster when using dynamic allocation?
Hard28What is the purpose of the 'Broadcast Variable' in the Spark architecture?
Medium29In a Spark application running on a Databricks cluster, the driver program creates a SparkSession and defines a series of transformations. When an action is triggered, the driver requests resources from the cluster manager. Which component is responsible for negotiating and acquiring these resources on behalf of the Spark application?
Easy30A Spark job is running on Databricks and experiences a stage where tasks are taking much longer than expected. The Spark UI shows that some tasks have significantly higher shuffle read sizes than others, and the stage is skewed. Which Spark feature can automatically mitigate this skew by splitting large partitions into smaller ones?
Hard31Refer to the exhibit. Based on the error log, what is the most likely cause of the job failure?
Hard32A developer is using Spark on Databricks and notices that a particular job has many stages due to shuffle operations. They want to understand the role of the shuffle in the Spark execution model. Which two statements accurately describe the behavior of a shuffle operation in Spark? (Choose two.)
Medium33A data engineer runs a PySpark job on a Databricks cluster. The job reads a 500 GB Parquet dataset, applies a filter, and writes the result. The engineer notices that during execution, all tasks of a particular stage complete quickly except for a handful that take far longer, and the Spark UI shows these tasks are processing partitions that contain far more records than others. Which Spark architecture concept best explains this behavior, and what is the most appropriate remediation?
Medium34A developer is debugging a Spark job and observes that a particular stage has 200 tasks, but only 10 executors with 2 cores each are available. What will happen to the remaining tasks in that stage?
Medium35A developer is tuning a Databricks job and wants to know how many tasks will be created for the final stage of a job that reads a Parquet file with 200 partitions, applies a filter, and then calls coalesce(10) before writing the result. Assuming no other repartitioning or shuffles occur, how many tasks will the final write stage contain?
Medium36Which TWO factors influence the effective parallelism of a Spark application?
Medium37What is the primary function of the Spark DAG Scheduler?
Medium38A data engineer is configuring a Spark application on Databricks. They set `spark.executor.instances` to 4, `spark.executor.cores` to 5, and `spark.executor.memory` to 16g. The cluster has 5 worker nodes, each with 16 cores and 64 GB RAM. What is the maximum number of tasks that can run concurrently across all executors?
Easy39A data engineer observes that a Spark Structured Streaming job on Databricks processes micro-batches with steadily increasing latency over several hours. The Spark UI shows that the number of active tasks per batch stays constant, but each task processes a growing amount of state. Which architectural behavior explains this pattern?
Hard40Which TWO factors contribute to the 'Data Locality' optimization in Spark?
Medium41Which THREE components are involved in the process of executing a Shuffle operation?
Medium42A developer is using Spark on Databricks and wants to understand how the Driver and Executors communicate during a job. Which two statements accurately describe this interaction? (Choose two.)
Medium43Which TWO of the following statements correctly describe the role of the Spark Executor in a Databricks environment?
Medium44Which Spark configuration property determines the maximum amount of memory the Spark Driver can request for itself when running on a Kubernetes cluster?
Medium45A Spark job reads a large Parquet dataset, performs a filter, and then a groupBy aggregation. The job's DAG shows two stages: one for the filter and one for the aggregation. The first stage has 200 tasks, and the second stage has 200 tasks. The job is running on a cluster with 10 executors, each with 8 cores. The engineer observes that the second stage takes significantly longer than the first. Which of the following is the most likely cause for the increased duration in the second stage?
Hard46A Databricks engineer is diagnosing why a Spark job's shuffle phase writes a very large amount of data to disk. The engineer wants to reduce shuffle overhead by changing how the job is structured and configured. Which TWO actions are most likely to reduce the volume of shuffle data written? (Choose two.)
Hard47In the Databricks Spark environment, what is the role of the 'Shuffle Service'?
Medium48Which component manages the lifecycle and allocation of executors in a Databricks cluster?
Easy49What happens when an action is called on a Spark DataFrame?
Easy50A Spark application is running in cluster mode on Databricks. The driver program is running on a worker node, and the application has been running for several hours. Suddenly, the driver node experiences a hardware failure and crashes. What happens to the running tasks and the application?
Hard51Which component in the Spark architecture is responsible for maintaining the state of the Spark application and coordinating the execution of tasks across the cluster?
Medium52Which property of RDDs (Resilient Distributed Datasets) is primarily responsible for Spark's fault tolerance during cluster execution?
Hard53A Spark application is running on a cluster with 5 executors. The driver program creates a broadcast variable that is used in a transformation. Which two components are directly involved in distributing and using the broadcast variable? (Choose two.)
Medium54A data engineer submits a Spark application using spark-submit in client deploy mode from an edge node. The application reads a large Parquet dataset, performs a groupBy aggregation, and writes the result to a Delta table. The engineer notices that the Driver process runs on the edge node and remains alive throughout the application's lifetime. Which statement best describes the role of the Driver in this scenario?
Medium55Which TWO of the following statements accurately describe the role of the Spark Executor in a cluster deployment?
Hard56What is the primary role of the 'Cluster Manager' in Spark?
Medium57A developer submits a Spark application to a Databricks cluster using spark-submit with deploy mode set to cluster. During execution, one of the worker nodes hosting a task fails and is lost by the cluster manager. Which Spark component is responsible for rescheduling the failed task on another available executor?
Easy58Which TWO of the following statements accurately describe the relationship between Spark Executors and memory management within a Databricks cluster?
Hard59What is the primary role of the 'Executor' process in the Spark distributed architecture?
Easy60Which component in the Spark cluster architecture is responsible for communicating directly with the Cluster Manager (e.g., YARN, Mesos, K8s) to request and release resources?
Easy61Which process is responsible for tracking the location of data blocks cached in the executors?
Medium62Which Spark component is responsible for maintaining the Directed Acyclic Graph (DAG) of stages and tasks?
Easy63In a Databricks Spark cluster, which component is primarily responsible for scheduling tasks and managing the distribution of computation across the worker nodes?
Medium64What is the consequence of having 'wide dependencies' in a Spark job regarding the Spark Architecture?
Hard65A data engineer runs a Spark job on a Databricks cluster using the default FIFO scheduler. They notice that a long-running job is holding all cluster resources, and short ad-hoc queries submitted later are stuck waiting. The engineer wants to allow concurrent scheduling of multiple jobs within the same Spark application so that short jobs can run while the long job is still executing. Which Spark configuration should be set to enable this behavior?
Medium66In the context of the Spark Driver, what is the 'DAG' and why is it important?
Easy67What is the function of the 'Executor' within the Spark execution model?
Medium68Refer to the exhibit. Which of the following is the most likely cause for this 'shuffle fetch failure' in a Databricks cluster?
Medium69Which of the following describes the 'Driver' process in a Spark application?
EasyOther domains
All Databricks-Spark-Assoc exam domains
Frequently asked questions
- What does the Spark Architecture and Components domain cover on the Databricks-Spark-Assoc exam?
- Be able to trace what happens from a DataFrame action to jobs, stages, and tasks, and identify which component does what. The key point: the driver builds the plan and schedules work, while executors run tasks and hold cached data.
- How many questions are in this domain?
- This page lists all 69 Spark Architecture and Components questions in the Databricks-Spark-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Spark Architecture and Components questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.