Courseiva

Databricks-Spark-Assoc · domain

scenario questions

Practise Databricks Certified Associate Developer for Apache Spark scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

295 questions66 easy142 medium87 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (295)

Click any question to see the full explanation, or start a practice session above.

1

In Spark's cluster architecture, what happens to the tasks if the Driver node crashes during the execution of a job?

Easy
2

A developer is writing a Spark SQL query that must handle null values in a column named discount. The requirement is to replace null discounts with 0.0, but also to replace any negative discount values with 0.0, leaving positive values unchanged. Which expression should be used?

Hard
3

Which component in Structured Streaming is responsible for providing fault tolerance and ensuring data is processed exactly once?

Easy
4

You have a Pandas-on-Spark DataFrame 'psdf'. You perform an operation that results in a 'compute.ops_on_diff_frames' error. What is the root cause of this behavior?

Hard
5

When performing a join between a very small table and a massive table, which Spark SQL optimization technique should be applied to prevent a full shuffle?

Medium
6

A developer is writing a Spark Connect application that must run a pandas UDF on a remote Databricks cluster. The code uses the spark.conf.set() API to pass a custom Python module to the executors. Which statement describes the correct behavior?

Medium
7

Which SQL command is used to view the history of operations performed on a Delta table, including timestamps and operation types?

Easy
8

A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?

Hard
9

A developer is using Spark SQL to process a streaming DataFrame from a Kafka source. They need to perform a stateful operation that maintains state across micro-batches to count occurrences of each key over a sliding window of 10 minutes, sliding every 5 minutes. Which Spark SQL operation should they use?

Hard
10

When performing a 'Z-ORDER' operation on a Delta table, how does it improve query performance?

Medium
11

You are optimizing a Spark SQL query that performs a large join between a 10GB table and a 5MB lookup table. To ensure performance efficiency, which command should you use to hint to the optimizer?

Medium
12

Refer to the exhibit. Which configuration issue is most likely causing this warning in a Spark application?

Hard
13

What is the primary purpose of the 'pyspark.pandas' module in the Databricks environment?

Easy
14

When analyzing a Spark job's execution plan, what does a 'BroadcastHashJoin' indicate compared to a 'SortMergeJoin'?

Medium
15

A developer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must tolerate occasional downstream failures and reprocess data without duplicates. The developer sets a checkpoint location and uses the default output mode. Which statement correctly describes how the checkpoint location contributes to fault tolerance in this scenario?

Medium
16

You are using Spark SQL to analyze a Delta table named transactions that is partitioned by a column region. You need to run a query that filters on region and also on a non-partitioned column amount. The table has statistics collected on region and amount. Which of the following best describes how Spark SQL will optimize the query?

Hard
17

A Spark application running on Databricks uses broadcast joins for a small dimension table. A developer notices that the broadcast variable is not being sent to executors as expected, causing a shuffle instead. Which configuration property directly controls the maximum size of a table that Spark will automatically broadcast?

Hard
18

A developer is writing a Spark Connect application that needs to read a CSV file from cloud storage and then perform a groupBy aggregation. The developer wants to minimize data transfer between the client and the server. Which of the following approaches best achieves this?

Medium
19

A data engineer has a PySpark DataFrame `events` with columns `event_ts` (TimestampType) and `user_id`. They must produce a new DataFrame where each row shows the event and the timestamp of that same user's previous event, ordered by `event_ts` within each `user_id`. Which code snippet correctly accomplishes this?

Medium
20

You are querying a large Delta table partitioned by 'event_date'. You need to calculate the daily count of unique users, but the query is running slowly due to data skew on high-traffic dates. Which SQL approach effectively mitigates this skew during the aggregation?

Medium
21

A developer is working with a Pandas API on Spark DataFrame `psdf` that has a default index. They call `psdf.sort_values('amount')` and then attempt to use `.loc` with an integer label to retrieve a specific row. They find that the integer label does not correspond to the row position they expect. What is the most likely explanation?

Hard
22

Which term describes the unit of work that is dispatched by the Driver to a specific Executor?

Easy
23

Which of the following describes the behavior of a Delta table when a 'DELETE' operation is performed?

Medium
24

A developer is writing a Spark Connect application that uses a Python UDF to transform a column. The UDF depends on a third-party Python library that is installed on the developer's laptop but not on the Databricks cluster. The developer runs the code and receives an error indicating the module cannot be found. What is the most appropriate fix?

Hard
25

A Spark application on Databricks uses a broadcast variable to distribute a large lookup table to all executors. The developer notices that the broadcast variable is not being used efficiently, as executors are still fetching the data multiple times. Which component is responsible for ensuring that the broadcast data is distributed only once per executor and cached there?

Medium
26

When a Spark application is running in Databricks, what determines the number of tasks that can run in parallel?

Medium
27

In the context of Databricks, what occurs when a Spark stage is described as 'Shuffle-heavy'?

Hard
28

A developer needs to add a column `full_name` to DataFrame `people` by concatenating `first_name` and `last_name` with a single space, and must handle rows where `last_name` is null by producing just the first name. Which expression is correct?

Medium
29

A developer is troubleshooting a Spark Connect client that intermittently fails with connection errors to a Databricks cluster. Which two configuration practices help ensure stable connectivity? (Choose two.)

Medium
30

A developer is building a Structured Streaming job that reads from a Delta table as a stream and writes to another Delta table. The job must support exactly-once processing and allow the output to be updated incrementally. Which two options are required to achieve exactly-once semantics? (Choose two.)

Medium
31

A data engineering team is migrating a client application to use Spark Connect. The application connects to a remote Databricks cluster. Which architectural component processes the client's DataFrame operations and executes them against the Spark cluster?

Medium
32

What is the purpose of the 'ANALYZE TABLE' command in Databricks?

Easy
33

A developer is migrating a legacy PySpark application to use Spark Connect. Which architectural change is fundamental to how Spark Connect executes operations compared to traditional Spark sessions?

Medium
34

A data engineer has a DataFrame `df` with columns `order_id`, `customer_id`, and `order_total`. They need to create a new DataFrame containing only the rows where `order_total` is greater than 100 and only the columns `order_id` and `order_total`. Which combination of DataFrame operations accomplishes this most efficiently?

Easy
35

Based on the exhibit, what is the most likely reason for this error in a Databricks notebook?

Medium
36

You are developing a data pipeline using Pandas API on Spark. You need to perform operations that are efficient and avoid unnecessary data shuffling. Which two of the following operations are considered expensive because they may trigger a full shuffle or collect data to the driver? (Choose two.)

Medium
37

A data engineer is tuning a Spark Structured Streaming job on Databricks that reads from a Kafka topic with 12 partitions. The job uses a static allocation of executors, each with 4 cores. The engineer notices that only 4 tasks are running concurrently, even though there are 12 Kafka partitions and 3 executors are available. Which Spark configuration is most likely causing this limitation?

Medium
38

A data engineer is using Pandas API on Spark and needs to perform a join between two Pandas-on-Spark DataFrames `psdf1` and `psdf2` on a common column `id`. They write the following code: ```python result = psdf1.merge(psdf2, on='id', how='inner') ``` Which statement best describes the execution and potential issue with this operation?

Medium
39

A Databricks notebook job calls a PySpark UDF built on a Python function that performs string parsing. The job completes successfully on a small sample, but on the full production dataset it fails with a PythonException and the executor logs show high garbage collection time. Which change is the most appropriate first step to make the job reliable without changing business logic?

Medium
40

A developer is using Pandas API on Spark and needs to perform operations that involve multiple DataFrames. They encounter a 'compute.ops_on_diff_frames' error. Which two actions can resolve this error? (Choose two.)

Medium
41

You are migrating a legacy Pandas codebase to Databricks using the Pandas API on Spark. You have a DataFrame 'pdf' and need to calculate the average of a column 'revenue' while ensuring the computation remains distributed across the cluster. Which command is the idiomatic approach to achieve this?

Medium
42

A Spark job is experiencing data skew during a join operation on a key column. Which strategy is most effective for mitigating this issue without changing the business logic?

Medium
43

Which library import is required to enable the Pandas API on Spark within a Databricks notebook?

Easy
44

A developer has a DataFrame `df` with columns `id` and `score`, and several rows contain null values in `score`. They want a new DataFrame in which rows with a null `score` are removed, keeping only rows where `score` is present. Which single call achieves this?

Easy
45

In the context of the Spark Driver, which component is specifically responsible for tracking the location of cached data blocks across the executors?

Medium
46

Which THREE factors should a developer consider when choosing a partition count for a shuffle operation?

Medium
47

A developer is writing a Structured Streaming query that reads from a JSON file source and writes to the console for debugging. The query uses `outputMode("append")`. Which statement describes the output behavior?

Easy
48

A developer has a DataFrame `df` and wants to remove duplicate rows considering only the columns `user_id` and `event_type`, keeping the first occurrence according to the current row order. Which DataFrame operation achieves this?

Easy
49

A data engineer must produce a report that shows each department's total salary, but only for departments where the total salary exceeds 500,000. The source DataFrame is created from a Delta table with columns department and salary. Which Spark SQL query correctly returns the desired result?

Medium
50

You have a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. You need to ensure that the job processes each record exactly once, even after failures. Which of the following should you configure?

Medium
51

An analytics engineer is designing a Structured Streaming pipeline that reads from an Apache Kafka topic and writes JSON-formatted output to cloud object storage. The pipeline must maintain exactly-once processing semantics and support automatic recovery from cluster restarts. Which TWO actions are mandatory to achieve these requirements? (Choose 2)

Hard
52

Which of the following best describes the purpose of a Spark Session in a Databricks environment?

Easy
53

Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?

Hard
54

A data engineer has a PySpark DataFrame `readings` with columns `sensor_id` (string) and `celsius` (double). The engineer must produce a new DataFrame where every temperature is converted to Fahrenheit using the formula `celsius * 9/5 + 32`, while keeping both the original `sensor_id` and a column named `fahrenheit`, and must avoid collecting data to the driver. Which DataFrame operation should be used?

Medium
55

Which SQL function is used to create a temporary view that persists only for the duration of the current Spark session?

Easy
56

A data engineer submits a PySpark job that performs a wide transformation via a join operation across two large datasets. During execution, several tasks in the shuffle stage fail repeatedly due to transient network timeouts between worker nodes. How does Apache Spark's architecture handle these failed tasks?

Medium
57

Which of the following is the most efficient way to convert a Spark DataFrame into a format suitable for low-latency SQL queries in Databricks?

Medium
58

A data analyst uses Spark Connect to run a PySpark job against a Databricks cluster. They call `df.show()` to preview the DataFrame. Where does the actual computation for `show()` occur?

Easy
59

A developer wants to start using Spark Connect from a local Python environment to connect to an existing Databricks cluster. Which step is required to establish the connection?

Easy
60

A data engineer has a Pandas-on-Spark DataFrame `psdf` with a default index. They call `psdf.sort_values('amount', ascending=False)`. After the operation, they notice that the resulting DataFrame's index values no longer match the original row positions. Which statement best describes the behavior of the index after sorting?

Medium
61

A Databricks Structured Streaming job reads JSON files from a cloud storage directory and writes aggregated results to a Delta table. The source directory receives new files continuously and files are never modified after being written. The pipeline must tolerate late-arriving event data by up to 30 minutes and must not reprocess already-emitted windows when the query is restarted. Which combination of configurations is required to meet these requirements?

Hard
62

A developer runs `spark.sql("SELECT * FROM sales")` and receives an error stating the table or view cannot be found, even though a Parquet directory exists at `dbfs:/mnt/raw/sales/`. The developer wants to query that Parquet data using Spark SQL without moving or copying the files. Which action should the developer take?

Easy
63

A data engineer is processing a 500 GB Parquet dataset on Databricks using the Pandas API on Spark. They need to extract a single scalar value, the maximum timestamp, to pass to a downstream orchestration tool. They use psdf['timestamp'].max(). Which statement correctly describes how this operation executes?

Medium
64

A developer runs a PySpark job on Databricks that filters a Delta table and then calls count() and show(). The Spark UI shows two separate scans of the same table. The developer wants to avoid the second scan without changing the filter logic. Which action should be taken?

Medium
65

A Spark application running on a Databricks cluster uses a broadcast variable to distribute a small lookup table to all executors. During execution, the driver serializes the broadcast variable and sends it to each executor. Which component is responsible for storing the broadcast data on the executor side and making it available to tasks?

Hard
66

A developer is working with a Spark SQL DataFrame that contains a column 'tags' which is an array of strings. The developer needs to filter rows where the array contains the string 'spark' and also transform the array to uppercase. Which TWO Spark SQL functions should be used to achieve these requirements? (Choose two.)

Medium
67

A developer is using the Pandas API on Spark and wants to write efficient code. They are reviewing operations that can cause a full shuffle of data across the cluster. Which TWO operations should they be cautious about because they typically require a shuffle? (Choose two.)

Hard
68

A developer is using Spark Connect to connect to a Databricks cluster. They attempt to read a CSV file from a path that exists on the client machine's local disk. The operation fails with a file not found error. What is the most likely reason for this failure?

Hard
69

Which THREE components are part of the Spark execution environment that resides on the Driver node?

Medium
70

A developer is using Spark SQL to join two DataFrames: a large fact table 'orders' and a smaller dimension table 'customers'. The join key is customer_id. The developer notices that the join is causing a shuffle and wants to avoid it. Which Spark SQL technique should be used to eliminate the shuffle for this join?

Hard
71

What is the primary benefit of using Spark Connect in a Databricks environment compared to traditional Spark clients?

Easy
72

Which TWO of the following are true concerning Spark SQL's handling of NULL values?

Hard
73

A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?

Hard
74

A developer is troubleshooting a Spark Connect client that intermittently fails to create a session against a Databricks cluster. The cluster is configured to auto-terminate after 20 minutes of inactivity. The client script runs on a schedule every hour. What is the most likely cause of the intermittent session creation failures?

Hard
75

A developer is tuning a Spark job that performs a join between a large fact table and a medium-sized dimension table. The job suffers from data skew, with a few keys having a disproportionately large number of rows. The developer wants to mitigate the skew. Which two actions are most effective? (Choose two.)

Hard
76

An engineer is developing a Structured Streaming job that reads from an Apache Kafka source and writes the output continuously toDelta Lake using outputMode("append"). The stream occasionally experiences late-arriving data. Which downstream behavior can the engineer expect regarding the Delta Lake table?

Medium
77

A developer writes a Spark Connect application that calls `df.cache()` on a large DataFrame, then performs several transformations and an action. The developer expects the cached data to persist on the client for reuse across sessions. Which statement describes what actually happens?

Medium
78

A data engineer needs to add a column `rank_in_dept` to a DataFrame `employees` that ranks each employee by `salary` descending within their `department`, but they must not collapse rows. They also want ties to receive the same rank with gaps afterward. Which expression correctly produces this column using the DataFrame API?

Hard
79

A Spark job reads a large Parquet file, performs a groupBy operation, and then writes the result. During execution, the job fails with an OutOfMemoryError on the Driver. Which component is most likely responsible for the memory issue?

Hard
80

A developer submits a Spark application to a Databricks cluster. The application creates a SparkSession, reads a CSV file, and calls count() on the resulting DataFrame. Which component is responsible for translating this logical operation into a physical execution plan and coordinating its execution across the cluster?

Easy
81

A developer runs a Spark application on a Databricks cluster in Standard access mode. The application reads a Parquet file, applies a filter, and calls `df.cache()` before an action. During execution, the driver logs show that a stage is retried because a task failed with an executor lost error. Which component is responsible for rescheduling the failed task on another executor within the same application?

Medium
82

A developer is writing a Spark application that will run on a Databricks cluster. They need to ensure that the driver program can communicate with the executors and that tasks are distributed correctly. Which component is responsible for coordinating the execution of tasks across the executors?

Easy
83

A developer wants to use the pandas API on Spark in a Databricks notebook. They have an existing PySpark DataFrame `sdf`. Which code snippet correctly creates a pandas-on-Spark DataFrame from `sdf` while preserving the distributed execution plan?

Easy
84

A data engineer has two Spark SQL DataFrames: customers (customer_id, name) and orders (order_id, customer_id, amount). They want to retrieve every customer along with their orders, but they also want to include customers who have placed no orders, showing null for the order columns. Which operation should they use?

Medium
85

A developer must join a 4 TB `transactions` DataFrame against a 900 MB `merchants` DataFrame on `merchant_id`. The cluster has 40 executors each with 16 GB of memory, and the job currently shuffles the large side. They want to avoid the shuffle entirely. Which change should they make?

Hard
86

A developer is writing a PySpark script that must run a SQL statement against DataFrames already registered as temporary views named orders and returns. The developer wants the query to use Spark SQL syntax while returning a DataFrame that can be further transformed with the DataFrame API. Which call accomplishes this?

Easy
87

Which of the following describes the function of the checkpoint location in a Structured Streaming query?

Easy
88

Refer to the exhibit. Which performance indicator suggests that Task 15 is likely causing a performance bottleneck during the execution of a join operation?

Hard
89

Which clause is used in a SELECT statement to filter the results based on aggregated values?

Easy
90

A data analyst wants to use Pandas API on Spark in a Databricks notebook but is unsure how to import it. Which import statement correctly enables the Pandas API on Spark?

Easy
91

A Spark job reads a large CSV file, performs a groupBy aggregation, and then writes the result. The Spark UI shows that the job has multiple stages, and one stage has a large number of tasks. Which factor primarily determines the number of tasks in the stage that performs the aggregation?

Medium
92

A developer is working with a Pandas API on Spark DataFrame `psdf` and wants to perform operations that are efficient in a distributed environment. Which two operations are considered efficient and do not require collecting data to the driver? (Choose two.)

Medium
93

You are processing JSON data from a stream. You need to extract nested fields from the JSON structure. Which function is the most efficient and standard way to handle this in Structured Streaming?

Medium
94

A developer is writing a Spark Connect application and wants to create a SparkSession that connects to a remote Databricks cluster. The developer has the cluster's connection string and an access token. Which method should be used to build the session?

Easy
95

A Spark application is submitted to a Databricks cluster. The application uses a broadcast variable to distribute a small lookup table to all Executors. Which component is responsible for broadcasting this variable?

Medium
96

Which configuration parameter should be adjusted to change the default number of partitions when reading from a shuffle-heavy operation?

Medium
97

You are building a Structured Streaming job that reads from a Kafka source and writes to a Delta table. You need to ensure that the job can recover from failures and process data exactly once. Which TWO of the following are required to achieve exactly-once semantics? (Choose two.)

Medium
98

A data engineer is using Pandas API on Spark to process a large dataset. They call `psdf.to_pandas()` on a DataFrame that is 50 GB in size. What is the most likely outcome?

Hard
99

Which of the following correctly describes the relationship between a Spark Job and a Spark Stage?

Medium
100

A developer is building a feature that must, for each `customer_id`, concatenate the distinct `product` values from many rows into a single comma-separated string. The result must contain each product only once per customer. Which single approach produces this result?

Medium
101

A developer has two DataFrames, `orders` (columns `order_id`, `customer_id`) and `customers` (columns `customer_id`, `customer_name`). They need a result containing every order, with the matching customer name where one exists and null where the customer is not found. Which single join configuration guarantees this?

Hard
102

A developer writes a Spark Connect client that creates a DataFrame, calls `df.collect()`, and then reuses the same DataFrame for a second `df.count()`. The cluster is remote. What happens on the second action?

Hard
103

A developer runs a PySpark job that joins a 10 GB DataFrame with a 50 MB lookup DataFrame. The job takes far longer than expected, and the Spark UI shows a SortMergeJoin with a large shuffle read and write for both sides. The developer wants the smallest change that most improves performance. Which action should the developer take?

Easy
104

A data engineer has a batch DataFrame `df` with a `status` string column and wants to keep only rows where `status` equals "active". The engineer wants the filter applied as early as possible in the plan and does not want a shuffle. Which operation best fits this requirement?

Easy
105

A Spark job reads a large Parquet dataset, performs a groupBy on a high-cardinality column, and writes the result to a Delta table. The job fails with a FetchFailedException on a particular executor. The Spark UI shows that the executor had sufficient memory but the shuffle fetch failed due to a connection reset. Which configuration change is most likely to resolve this issue?

Hard
106

A data analyst is using the Pandas API on Spark to compute summary statistics. They call `psdf.describe()` on a large DataFrame and notice the job takes much longer than expected. They want to understand why this operation is more expensive than a similar operation on a small local pandas DataFrame. What is the primary reason?

Medium
107

A developer is building a Python application that connects to a Databricks cluster using Spark Connect. The application uses the `databricks-connect` package and is configured with the cluster ID and authentication credentials. During a test run, the developer calls `spark.sql("SELECT * FROM sales")` and then `df.show()`. What happens when the `show()` action is executed?

Medium
108

You are using Pandas API on Spark to process a large dataset. You have a Pandas-on-Spark DataFrame `psdf` and you apply a custom Python function using `psdf.apply(func, axis=1)`. The function is computationally intensive and you notice that the job is running slowly with many tasks. What is the most likely reason for the performance issue?

Hard
109

A streaming DataFrame job on Databricks writes to a Delta table every 10 seconds. Over several hours, the number of files in the target directory grows into the hundreds of thousands, and downstream reads slow dramatically. The job uses foreachBatch with a write that produces many small files per micro-batch. Which action should be taken to reduce the small-files problem for this streaming write?

Medium
110

You are processing a large, highly skewed PySpark DataFrame in Databricks and want to optimize a forthcoming join operation against a small lookup dimension table. Which TWO strategies are valid and effective DataFrame API techniques to optimize this join performance? (Choose TWO)

Hard
111

A developer needs to join `orders` (large) with `customers` (small, a few thousand rows) on `customer_id`. They want to broadcast the small table and confirm the broadcast actually took effect. Which combination of actions is correct?

Hard
112

A developer has a DataFrame `events` with columns `user_id`, `event_type`, and `payload` (a JSON string). They need to extract the `device` field from `payload` into a new column without changing the other columns or the row count. Which approach is correct?

Medium
113

You are processing a streaming dataset of sensor readings. You need to calculate the average temperature every 10 minutes, allowing data to arrive up to 2 minutes late. Which windowing approach correctly handles this requirement in Structured Streaming?

Medium
114

A data engineer wants to run a Structured Streaming query that reads from a Kafka topic and writes aggregated counts to a console sink for debugging. The query uses a grouping aggregation on a tumbling event-time window. Which output mode must be used so that only rows that changed since the last trigger are emitted?

Easy
115

A developer is building a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The job must handle late data up to 15 minutes and ensure that aggregations are updated correctly. The developer adds a watermark of 15 minutes on the event time column. What is the effect of this watermark on the aggregation state and output?

Hard
116

Which component in the Spark architecture is responsible for scheduling tasks and managing the execution of jobs on the cluster?

Medium
117

Which file format is best suited for performance-critical Spark applications that require efficient schema enforcement and column pruning?

Medium
118

You are developing a Structured Streaming job that reads from a Delta table and writes to another Delta table. You need to ensure that the streaming query can recover from failures and continue processing without data loss or duplication. Which of the following must be configured?

Medium
119

What happens when a Spark job triggers a 'shuffle' operation during execution?

Medium
120

Refer to the exhibit. What is the best way to resolve this error?

Hard
121

A Spark application is running on a Databricks cluster with 3 worker nodes, each having 4 cores. The application uses the default configuration. How many tasks can run concurrently across the cluster?

Easy
122

A developer is using Spark on Databricks and wants to monitor the progress of a job. They need to understand how the driver coordinates with executors. Which component is responsible for scheduling tasks onto executors and tracking their status?

Easy
123

A developer notices that a PySpark DataFrame transformation chain runs a full scan of a Delta table each time a new action is invoked, even though the source data has not changed. The developer wants to persist the intermediate DataFrame in memory across actions. Which method should be used?

Easy
124

Which action should be taken to optimize a Spark application that performs multiple operations on the same DataFrame and shows evidence of redundant re-computations in the Spark UI DAG visualization?

Easy
125

A developer has a DataFrame `raw` and needs to permanently persist it as Parquet partitioned by `region`, overwriting any existing data at that path, without registering it in the metastore. Which call achieves this?

Easy
126

What is the primary benefit of the Catalyst Optimizer in the Spark SQL architecture?

Medium
127

Which TWO of the following techniques effectively reduce the shuffle volume in a Databricks Spark job?

Hard
128

A data engineer is using Spark Connect to run a job on a Databricks cluster. They notice that when they call `df.count()`, the operation takes longer than expected. They suspect that the client is transferring data unnecessarily. Which statement best explains the data transfer behavior of `df.count()` in Spark Connect?

Hard
129

A data scientist is writing a Spark Connect application that requires custom user-defined functions (UDFs). How are UDFs handled when executing code through Spark Connect?

Hard
130

What is the primary function of the 'Shuffle Service' in a Spark cluster when using dynamic allocation?

Hard
131

What is the purpose of the 'Broadcast Variable' in the Spark architecture?

Medium
132

A developer is building a Spark Connect application that runs on a laptop and connects to a remote Databricks cluster. During development, the laptop loses network connectivity for a few minutes while a long-running DataFrame transformation is executing. The developer notices the local Python process raises a gRPC error and the job is no longer tracked. Which statement best explains this behavior?

Medium
133

Which state management strategy should you implement if you notice your streaming aggregation query is failing due to excessive memory consumption on the executor nodes?

Hard
134

A developer is writing a Spark Connect application that runs on a local laptop and connects to a Databricks cluster. The application defines a Python function and registers it with `spark.udf.register` for use inside a `select` expression. When the code runs, the function executes on the server. Which statement describes how the UDF is handled in this scenario?

Medium
135

In a Spark application running on a Databricks cluster, the driver program creates a SparkSession and defines a series of transformations. When an action is triggered, the driver requests resources from the cluster manager. Which component is responsible for negotiating and acquiring these resources on behalf of the Spark application?

Easy
136

A data engineer is writing a PySpark job that must read Parquet data, apply transformations, and then write the result. They want to avoid recomputation if the resulting DataFrame is referenced multiple times in later stages. Which action best accomplishes this?

Medium
137

You are designing a Structured Streaming job that must read from a file source and write to a Delta table. You want the job to be resilient to failures and to continue processing only new files after a restart. Which two actions should you take? (Choose two.)

Medium
138

An enterprise data engineering team is migrating legacy PySpark client applications to use Spark Connect to improve client stability and isolate resource consumption. A developer initializes the Spark session pointing to a remote cluster. Which specific mechanism does Spark Connect use to communicate execution plans between the client application and the server cluster?

Medium
139

A data engineer has a DataFrame `orders` with columns `order_id`, `customer_id`, and `amount`, and a small lookup DataFrame `tiers` with `customer_id` and `tier`. The engineer wants to attach the tier to every order. Some orders have a customer_id that is not present in tiers, and those orders must still appear with a null tier. Which operation produces this result?

Hard
140

A data engineer runs the following statement in a Databricks notebook: CREATE OR REPLACE TEMP VIEW high_value_customers AS SELECT customer_id, SUM(amount) AS total FROM sales GROUP BY customer_id HAVING SUM(amount) > 10000. Later, the same engineer opens a new notebook attached to the same cluster and tries to run SELECT * FROM high_value_customers. What will happen?

Medium
141

A data engineer is writing a Spark SQL query that joins a `transactions` table to a `customers` table on `customer_id`. The engineer wants to ensure that rows from `transactions` with no matching customer are still returned, with nulls for customer columns, and also wants to exclude duplicate rows that arise from the join. Which TWO clauses should the engineer include? (Choose two.)

Medium
142

A data engineer is using Spark Connect from a remote Python client to interact with a Databricks cluster. The engineer wants to understand which operations are executed on the server side versus the client side. Which two statements correctly describe this behavior? (Choose two.)

Medium
143

A Spark job is running on Databricks and experiences a stage where tasks are taking much longer than expected. The Spark UI shows that some tasks have significantly higher shuffle read sizes than others, and the stage is skewed. Which Spark feature can automatically mitigate this skew by splitting large partitions into smaller ones?

Hard
144

Which clause is used in a Spark SQL query to limit the number of rows returned by a query, and in which logical order is it executed?

Easy
145

A Databricks job fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver for local processing. Which action should you take to resolve this?

Easy
146

When a Spark job is stuck in a shuffle phase, what is the most effective first step to identify the root cause of the performance bottleneck?

Hard
147

A streaming DataFrame `df` has a watermark defined on `eventTime` with a delay of 5 minutes. The query uses `withWatermark("eventTime", "5 minutes")` and writes to a Delta table in `append` output mode. A record with event time 10:00 arrives when the watermark is at 10:10. What happens to this record?

Hard
148

A streaming job using 'mapGroupsWithState' is failing due to excessive memory usage. Which strategy is most effective for mitigating this?

Hard
149

Refer to the exhibit. Based on the error log, what is the most likely cause of the job failure?

Hard
150

A developer is using Pandas API on Spark in a Databricks notebook and needs to combine two Pandas-on-Spark DataFrames that originate from different Spark DataFrame ancestors. They encounter a `compute.ops_on_diff_frames` error. Which two actions will resolve this error? (Choose two.)

Hard
151

You are analyzing a large dataset using Pandas API on Spark. You have a Pandas-on-Spark DataFrame `psdf` that was created from a Spark DataFrame with multiple partitions. You call `psdf.head(10)` to quickly inspect the data. What does this operation return?

Medium
152

A Spark application reads a large CSV file and performs a series of transformations, including a filter and a join with a small lookup table. The job is running slowly, and the Spark UI shows that the CSV parsing stage is taking a long time. Which action would most improve performance?

Medium
153

A developer is using Spark on Databricks and notices that a particular job has many stages due to shuffle operations. They want to understand the role of the shuffle in the Spark execution model. Which two statements accurately describe the behavior of a shuffle operation in Spark? (Choose two.)

Medium
154

What is the primary difference between a micro-batch streaming query and a continuous processing query in Spark?

Easy
155

You are monitoring a Structured Streaming query in Databricks and want to see the current status, including the number of input rows per second and the batch duration. Which of the following is the most direct way to access this information?

Easy
156

You are writing a Structured Streaming query that reads from a Kafka topic and outputs to the console. You want to see only the newly arrived data in each micro-batch, without aggregations. Which output mode should you use?

Easy
157

A data engineer runs a PySpark job on a Databricks cluster. The job reads a 500 GB Parquet dataset, applies a filter, and writes the result. The engineer notices that during execution, all tasks of a particular stage complete quickly except for a handful that take far longer, and the Spark UI shows these tasks are processing partitions that contain far more records than others. Which Spark architecture concept best explains this behavior, and what is the most appropriate remediation?

Medium
158

A developer has a PySpark DataFrame `df` with columns `order_id`, `customer_id`, and `order_ts` (timestamp). They need to return only the most recent order per customer, keeping all original columns, and they want to avoid a self-join or a manual sort-then-dropDuplicates approach. Which DataFrame operation should they use?

Medium
159

A developer is debugging a Spark job and observes that a particular stage has 200 tasks, but only 10 executors with 2 cores each are available. What will happen to the remaining tasks in that stage?

Medium
160

A data engineer has a PySpark DataFrame `events` with columns `user_id`, `event_time` (timestamp), and `payload` (string). They must produce a new DataFrame where each row is enriched with the `payload` value from the user's immediately preceding event, ordered by `event_time`, without collapsing any rows. Which TWO approaches accomplish this? (Choose two.)

Medium
161

You need to combine two datasets in Spark SQL, retaining all records from the left table and matching records from the right table, while filling unmatched right columns with null values. Which join type should you use?

Medium
162

A developer is using Pandas API on Spark and wants to convert a Pandas-on-Spark DataFrame `psdf` back to a standard pandas DataFrame for local analysis. Which method should they use?

Easy
163

A developer is tuning a Databricks job and wants to know how many tasks will be created for the final stage of a job that reads a Parquet file with 200 partitions, applies a filter, and then calls coalesce(10) before writing the result. Assuming no other repartitioning or shuffles occur, how many tasks will the final write stage contain?

Medium
164

A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?

Hard
165

You are using Spark SQL to join two large Delta tables, orders and customers, on a common column customer_id. The orders table is partitioned by order_date, and the customers table is not partitioned. You need to ensure the join is efficient and minimizes shuffling. Which TWO actions should you take? (Choose two.)

Hard
166

In Spark SQL, what is the primary difference between a temporary view and a global temporary view?

Easy
167

Which TWO factors influence the effective parallelism of a Spark application?

Medium
168

Which environment variable is mandatory to establish a connection to a Databricks cluster using Spark Connect in a local Python environment?

Easy
169

A structured streaming DataFrame writes to a Delta table with a foreachBatch function that performs an upsert. After a cluster restart, the stream reprocesses some micro-batches and duplicate rows appear in the target table. The foreachBatch code already uses MERGE keyed on a unique id. Which change best prevents duplicates after restart?

Hard
170

You are processing a large dataset in Spark SQL and need to ensure that small files are avoided when writing data to Delta Lake. Which approach effectively minimizes small file generation during write operations?

Medium
171

A developer notices that a DataFrame transformation chain is executed twice: once for a count action used for logging and again for a write action. The source is a large Delta table and the repeated scan adds several minutes. Which action avoids the duplicate computation with the least risk?

Easy
172

An analyst notices that a Spark SQL query filtering a Delta table with `WHERE order_date = '2024-03-15'` scans far more data than expected, even though the table is partitioned by `order_date`. The partition column was loaded as a string in `yyyy-MM-dd` format. Which explanation best accounts for the excessive scan?

Medium
173

When joining two large tables in Spark SQL, which join type helps avoid expensive shuffles by loading a small table into memory on all executor nodes?

Medium
174

A developer uses a Spark Connect session to create a temporary view with `df.createOrReplaceTempView("sales_v")` and then runs `spark.sql("SELECT * FROM sales_v")` in the same session. What is the scope of that temporary view?

Hard
175

A developer wants to start a Structured Streaming query that reads from a Delta table and writes to another Delta table, and needs the query to process all existing data in the source table on its first run. Which option should be set on the read stream?

Easy
176

A developer runs a Spark Connect client session against a Databricks cluster with `spark.conf.set("spark.sql.shuffle.partitions", "400")`. The cluster is configured with 8 worker nodes. Which component actually applies the shuffle partition setting to the physical plan?

Easy
177

What is the primary function of the Spark DAG Scheduler?

Medium
178

A data engineer is configuring a Spark application on Databricks. They set `spark.executor.instances` to 4, `spark.executor.cores` to 5, and `spark.executor.memory` to 16g. The cluster has 5 worker nodes, each with 16 cores and 64 GB RAM. What is the maximum number of tasks that can run concurrently across all executors?

Easy
179

A data engineer observes that a Spark Structured Streaming job on Databricks processes micro-batches with steadily increasing latency over several hours. The Spark UI shows that the number of active tasks per batch stays constant, but each task processes a growing amount of state. Which architectural behavior explains this pattern?

Hard
180

A Databricks job joins a 500 GB sales table with a 300 GB returns table on a customer_id key. A few customer_id values account for a large fraction of rows on both sides, and the job fails with executor OOM during the join. The developer wants to distribute the hot keys across more partitions without changing the query logic. Which technique should be used?

Hard
181

A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)

Medium
182

When using Spark Connect, how does the client handle the authentication process with the Databricks workspace?

Medium
183

Which TWO factors contribute to the 'Data Locality' optimization in Spark?

Medium
184

An engineering team wants to execute PySpark queries locally from an Integrated Development Environment (IDE) while offloading all distributed compute and data processing to a remote Databricks cluster. Which Spark Connect component architecture makes this workflow possible?

Medium
185

Which THREE components are involved in the process of executing a Shuffle operation?

Medium
186

What is the primary purpose of the 'Cache' command in Spark SQL?

Medium
187

A developer is using Spark on Databricks and wants to understand how the Driver and Executors communicate during a job. Which two statements accurately describe this interaction? (Choose two.)

Medium
188

A developer is building a Spark SQL pipeline in a Databricks notebook. They need to persist an intermediate DataFrame, built from a transformation of a Delta table, as a physical table in the current database so other notebooks in the same cluster can query it. They also want the table metadata to be managed by the metastore and the data to reside in the default warehouse directory. Which Spark SQL statement should they use?

Medium
189

You have a Pandas API on Spark DataFrame `psdf` that was created from a Spark DataFrame with 200 partitions. You call `psdf.head(10)` in a Databricks notebook. What is the most likely performance characteristic of this operation?

Medium
190

A developer is using Spark Connect to connect to a Databricks cluster from a remote Python client. They need to run a custom Python function on a DataFrame column. They define the function and register it as a UDF using spark.udf.register(). After executing the job, they notice that the UDF fails with a ModuleNotFoundError for a library that is installed on their local machine but not on the cluster. What is the most likely cause and the appropriate solution?

Hard
191

A data scientist is working with a pandas-on-Spark DataFrame psdf that has a column 'category' with many unique values. They want to apply a custom Python function to each group to compute a complex statistic. They consider using psdf.groupby('category').apply(my_func). Which statement accurately describes the execution and potential performance implications of this operation?

Medium
192

A Databricks job reads a Parquet dataset, applies a chain of `withColumn` transformations, and writes it back with `.write.mode("overwrite").parquet(path)`. The Spark UI shows 6000 small output files totaling 50 GB, and a downstream reader is slow because of per-file overhead. You want to reduce the file count while keeping the write as a Spark-native operation. What should you do?

Medium
193

Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?

Medium
194

Which TWO of the following statements correctly describe the role of the Spark Executor in a Databricks environment?

Medium
195

A developer is building a Structured Streaming job in PySpark that reads from a Kafka topic and writes to a Delta Lake table. The job uses `outputMode("append")` and a 10-minute watermark on the event-time column `event_time`. A batch of late data arrives with events whose `event_time` is older than the watermark. What happens to these late events?

Medium
196

A developer writes a Spark SQL query that groups orders by region and computes the total revenue per region, but also needs to return the number of distinct customers per region in the same result set. Which TWO expressions correctly compute the distinct customer count per region in a single GROUP BY region query? (Choose two.)

Medium
197

Refer to the exhibit. Which action is the most likely cause of the error shown in the Spark job logs?

Medium
198

A developer is using Pandas API on Spark and encounters a `compute.ops_on_diff_frames` error when combining two Pandas-on-Spark DataFrames. Which two actions can resolve this error? (Choose two.)

Hard
199

An analyst has a Spark SQL DataFrame named events with a string column event_time in the format 'yyyy-MM-dd HH:mm:ss'. They want to add a new column event_date containing only the date portion, keeping the original column intact. Which expression should they use in a select statement?

Easy
200

A developer needs to add a computed column `discounted_price` equal to `price * 0.9` to an existing Delta table `products` and persist the change so all future queries see the new column. The table already contains data. Which statement should the developer run?

Hard
201

A job is reading a huge amount of data from a table, but only uses three columns. Which optimization technique will provide the most significant I/O performance benefit?

Medium
202

A developer has a local Python script that connects to a Databricks cluster using Spark Connect and creates a DataFrame from a small list of tuples. They then call .collect() on the DataFrame and receive the results. Which statement accurately describes how the data and operations are processed in this scenario?

Medium
203

Which Spark configuration property determines the maximum amount of memory the Spark Driver can request for itself when running on a Kubernetes cluster?

Medium
204

Which of the following Spark SQL configuration settings should be adjusted to prevent the 'Driver OOM' error when collecting massive amounts of query results to the driver node?

Medium
205

A data analyst needs to run a Spark SQL query that returns the top 5 highest-paid employees from a table named employees, ordered by salary descending. Which query correctly returns exactly five rows?

Easy
206

Which SQL function is used to concatenate strings while allowing you to specify a custom separator, handling null values by ignoring them?

Easy
207

A Spark job reads a large Parquet dataset, performs a filter, and then a groupBy aggregation. The job's DAG shows two stages: one for the filter and one for the aggregation. The first stage has 200 tasks, and the second stage has 200 tasks. The job is running on a cluster with 10 executors, each with 8 cores. The engineer observes that the second stage takes significantly longer than the first. Which of the following is the most likely cause for the increased duration in the second stage?

Hard
208

A Databricks engineer is diagnosing why a Spark job's shuffle phase writes a very large amount of data to disk. The engineer wants to reduce shuffle overhead by changing how the job is structured and configured. Which TWO actions are most likely to reduce the volume of shuffle data written? (Choose two.)

Hard
209

When executing a Spark SQL query, what does the Catalyst optimizer perform during the 'Analysis' phase?

Hard
210

In the Databricks Spark environment, what is the role of the 'Shuffle Service'?

Medium
211

Which component manages the lifecycle and allocation of executors in a Databricks cluster?

Easy
212

A developer has a Delta table events with a high-cardinality column user_id and a low-cardinality column country. A query filters on country = 'US' and also on user_id IN (...). The developer runs EXPLAIN and sees a full scan of all files. Which statement about data skipping and the Delta table's statistics correctly explains why the filter on country is not skipping files?

Hard
213

You are building a Structured Streaming pipeline that reads from a Delta table source and applies a stateful deduplication using dropDuplicates on a composite key. After several hours, the job fails with an error indicating that the state store has grown too large. You need to bound the state size while still removing duplicate events that arrive within a reasonable window. Which approach should you take?

Hard
214

Which THREE of the following are valid ways to monitor or debug Spark SQL query performance in Databricks?

Hard
215

A data engineer is working with the Pandas API on Spark and needs to convert a Spark DataFrame named `sdf` into a pandas DataFrame so it can be processed locally on the driver node. Which method should the engineer use to execute this conversion?

Medium
216

What happens when an action is called on a Spark DataFrame?

Easy
217

A Spark application is running in cluster mode on Databricks. The driver program is running on a worker node, and the application has been running for several hours. Suddenly, the driver node experiences a hardware failure and crashes. What happens to the running tasks and the application?

Hard
218

Which THREE of the following are valid ways to create a DataFrame from an existing table in Spark SQL?

Hard
219

Which component in the Spark architecture is responsible for maintaining the state of the Spark application and coordinating the execution of tasks across the cluster?

Medium
220

Which of the following describes the behavior of a 'Broadcast Hash Join' in Spark SQL?

Medium
221

A data engineer is building a PySpark application that must validate incoming records in a DataFrame `raw` before loading them into a curated table. They want to apply user-defined validation logic that cannot be expressed with built-in functions, and they want the result to remain a DataFrame column of Boolean values. Which TWO approaches allow this? (Choose two.)

Hard
222

A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes it locally. The developer wants to avoid the driver OOM while still obtaining the results. Which approach is most appropriate?

Hard
223

What is the result of applying the COALESCE function in Spark SQL when multiple arguments are provided?

Easy
224

A Spark job performing a join between a 10 GB table and a 5 MB lookup table is running slowly, and the physical plan shows a SortMergeJoin. You want to avoid the shuffle. What should you do?

Hard
225

Which property of RDDs (Resilient Distributed Datasets) is primarily responsible for Spark's fault tolerance during cluster execution?

Hard
226

A data engineer has a Pandas-on-Spark DataFrame `psdf` with a column `event_time` stored as string. They run `psdf['event_time'] = pd.to_datetime(psdf['event_time'])` where `pd` is the Pandas API on Spark module. What is the most likely outcome?

Medium
227

A Spark application is running on a cluster with 5 executors. The driver program creates a broadcast variable that is used in a transformation. Which two components are directly involved in distributing and using the broadcast variable? (Choose two.)

Medium
228

A streaming query reads from a rate source and writes to a Delta table using `outputMode("append")`. The developer observes that the query processes data continuously but the Delta table remains empty. Which condition explains why no rows are written?

Hard
229

A developer is using Structured Streaming with a Kafka source and wants to ensure that each message is processed exactly once, even in the event of failures. The developer has set a checkpoint location and is using `foreachBatch` to write to an external database. Which additional step is necessary to achieve exactly-once semantics?

Hard
230

A data engineer is building a feature pipeline and needs to add a monotonically increasing integer column `row_num` to a DataFrame `df` that assigns consecutive numbers to rows within each partition of a specified ordering, similar to a window function. Which approach uses the DataFrame API to compute this value?

Medium
231

Refer to the exhibit. Traceback (most recent call last): File "app.py", line 12, in <module> df = spark.read.table("default.sales") File "/opt/spark/python/pyspark/sql/session.py", line 314, in table return DataFrame(self._client.execute_plan(parser.parse_table(name)))) File "/opt/spark/python/pyspark/sql/connect/client/core.py", line 112, in execute_plan(y+"sessionID"), grpc.RpcError: StatusCode.UNAVAILABLE An engineer attempts to run a PySpark script using Spark Connect but encounters the traceback shown above. What is the most likely root cause of this execution failure?

Medium
232

A developer is using the Pandas API on Spark to process a large dataset. They need to apply a custom Python function to each value in a column. They consider using `psdf['col'].apply(custom_func)`. What should they be aware of regarding performance?

Medium
233

You are working with a Pandas-on-Spark DataFrame `psdf` that has a default index generated by Spark. You need to perform a join with another Pandas-on-Spark DataFrame `other` that also has a default index. After the join, you notice that the resulting DataFrame has a new index and the original indices are lost. Which of the following best explains this behavior?

Hard
234

A developer runs a PySpark job on Databricks that reads a large Delta table, filters on a timestamp column, and writes results to another Delta table. The job takes 45 minutes, but the Spark UI shows that 90% of task time is spent reading from the source table. The developer wants to reduce the read time. Which action should the developer take?

Medium
235

A data engineer submits a Spark application using spark-submit in client deploy mode from an edge node. The application reads a large Parquet dataset, performs a groupBy aggregation, and writes the result to a Delta table. The engineer notices that the Driver process runs on the edge node and remains alive throughout the application's lifetime. Which statement best describes the role of the Driver in this scenario?

Medium
236

When working with Delta Lake tables in Databricks, which command should you use to optimize the physical layout of files to improve query performance?

Medium
237

A developer notices a Spark job is failing with an OutOfMemoryError during a join operation on two large tables. The join key is highly skewed, causing one task to process significantly more data than others. Which technique should be applied to resolve this skew?

Medium
238

You are writing a Databricks notebook and want to use the Pandas API on Spark. Which import statement should you use to access the Pandas API on Spark?

Easy
239

A developer writes a Structured Streaming query that reads from a Kafka topic with `spark.readStream.format("kafka")` and then calls `.writeStream.format("console").start()`. The query runs, but after a few minutes the driver logs show that the query is only processing newly arriving offsets and older messages in the topic are never read. What is the most likely cause?

Medium
240

Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize performance?

Medium
241

A data engineer has a PySpark DataFrame `orders` with a string column `order_ts` formatted as `yyyy-MM-dd HH:mm:ss`. They need a new column `order_date` containing only the date portion as a `date` type, so downstream code can filter by day. Which approach is correct?

Easy
242

A developer is building a Spark Connect client application in Python that runs on a local workstation and connects to a remote Databricks cluster. The application must construct a DataFrame from a list of Python dictionaries without requiring the data to be uploaded to cloud storage first. Which approach should the developer use?

Medium
243

A data engineer has a DataFrame `events` with columns `user_id` and `ts`. They need to add a column `prev_ts` that holds the previous event timestamp for each user, ordered by `ts` ascending, without collapsing rows. Which operation accomplishes this?

Hard
244

A developer has a DataFrame `trades` with columns `trade_id`, `symbol`, and `price`. They want to add a column `prev_price` containing the price of the immediately preceding trade for the same symbol, ordered by `trade_id` ascending, without collapsing rows. Which transformation should they use?

Hard
245

Which TWO of the following statements accurately describe the role of the Spark Executor in a cluster deployment?

Hard
246

A job joining a large fact table with a small dimension table runs out of memory on executors during the join. The dimension table is about 40 MB after filtering and the configured spark.sql.autoBroadcastJoinThreshold is 10 MB. The join key is highly skewed in the fact table. Which action is most appropriate?

Hard
247

A developer is writing a PySpark job that must read a Parquet dataset, apply several transformations, and then write the result back to storage. They want to ensure the schema of the written data is inferred directly from the DataFrame rather than from any external definition, and they want to append to an existing Parquet directory. Which write configuration accomplishes this?

Medium
248

A data engineer is working with a Spark SQL DataFrame in Databricks that has a column named event_time stored as a string in the format 'yyyy-MM-dd HH:mm:ss'. They need to filter rows where event_time falls within the last 7 days relative to the current timestamp. Which Spark SQL expression correctly achieves this?

Medium
249

What is the primary role of the 'Cluster Manager' in Spark?

Medium
250

A developer submits a Spark application to a Databricks cluster using spark-submit with deploy mode set to cluster. During execution, one of the worker nodes hosting a task fails and is lost by the cluster manager. Which Spark component is responsible for rescheduling the failed task on another available executor?

Easy
251

A data analyst wants to connect a local Python script to a Databricks cluster using Spark Connect. The workspace URL and a personal access token are available. Which client-side step is required to create the remote Spark session?

Easy
252

A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?

Hard
253

Which TWO of the following statements accurately describe the relationship between Spark Executors and memory management within a Databricks cluster?

Hard
254

A Spark job writes a large DataFrame to a Delta table partitioned by date. The job is taking much longer than expected, and the Spark UI shows that many tasks are writing very small files. You have already set spark.sql.shuffle.partitions to 200. What is the most effective way to reduce the number of small files written?

Hard
255

A data engineer is building a Spark Connect application and wants to attach a small lookup table to every task without a shuffle. The table is 20 MB and the cluster has default settings. Which approach should the engineer use?

Medium
256

You are performing a stream-stream join between two streaming DataFrames, `orders` and `payments`, both with watermarks defined on their event time columns. The join condition is `orders.orderId == payments.orderId` and it is an inner join. What happens to state in this join?

Hard
257

You need to add a new calculated column 'discounted_price' to an existing DataFrame 'df' by multiplying 'price' by 0.9. Which DataFrame transformation accomplishes this correctly?

Easy
258

A developer wants to start a Structured Streaming query that reads from a Kafka topic and writes to the console for debugging. The developer uses `writeStream.format("console").start()`. What is the default trigger for this query?

Easy
259

You are troubleshooting a Spark Connect application that fails to connect to a Databricks cluster. The error message indicates an authentication failure. Which of the following is the most likely cause?

Medium
260

Which THREE of the following sources support streaming read operations in Spark Structured Streaming?

Medium
261

A developer is using the pandas API on Spark in a Databricks notebook. They have a pandas-on-Spark DataFrame psdf with a default index. They call psdf.sort_values('amount') and then psdf.head(10). They observe that the resulting index values are not sequential from 0 to 9, but instead appear as arbitrary integers. What is the most likely explanation for this behavior?

Hard
262

A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?

Medium
263

A data engineer needs to join two large DataFrames, `sales` and `products`, on the `product_id` column. The `products` DataFrame is extremely small and fits entirely in a single executor's memory. To optimize performance and avoid a costly shuffle join across the network, which strategy should be applied using the Spark DataFrame API?

Medium
264

A data engineer is using Spark Connect from a local Python environment to connect to a Databricks cluster. They attempt to use the spark.sparkContext.broadcast() method to broadcast a large lookup dictionary for use in a UDF. The code fails. What is the most likely reason for this failure?

Hard
265

A data engineer has two DataFrames: `orders` with columns `order_id` and `customer_id`, and `customers` with columns `customer_id` and `region`. They must produce a result containing only orders whose `customer_id` exists in `customers`, keeping every matching order row exactly once with no customer columns added. Which operation should be used?

Hard
266

A developer is using Spark Connect to run a PySpark application against a remote Databricks cluster. The application calls df.cache() on a DataFrame that is used multiple times. Which statement accurately describes how caching behaves in this scenario?

Medium
267

When using 'apply_batch' in Pandas-on-Spark, how does the function behave regarding the input data?

Hard
268

A developer is building a PySpark job that reads a Parquet dataset with 2,000 files into `df`. They call `df.cache()` and then execute three separate actions in the same session. The Spark UI shows the Parquet files are read from storage three times, and the cache never appears in the Storage tab. What is the most likely cause?

Hard
269

What is the primary role of the 'Executor' process in the Spark distributed architecture?

Easy
270

Which of the following describes the behavior of a 'Left Outer Join' in Spark SQL?

Easy
271

A data scientist is using Pandas API on Spark to process a large dataset. They call `psdf.apply(lambda row: row['a'] + row['b'], axis=1)` and notice extremely slow performance. Which statement best explains why this operation is inefficient and what alternative should be used?

Medium
272

You have a large DataFrame containing user transaction logs. You need to read this data and immediately repartition it by user_id to optimize downstream filtering operations. Which DataFrame API method should you use?

Medium
273

Which component in the Spark cluster architecture is responsible for communicating directly with the Cluster Manager (e.g., YARN, Mesos, K8s) to request and release resources?

Easy
274

A developer must combine two DataFrames, `left_df` and `right_df`, on a key column `id`. They need every row from `left_df` regardless of whether a match exists in `right_df`, and matching rows from `right_df` where available, with unmatched right-side columns filled as null. Which join invocation produces this result?

Hard
275

A developer has a Spark SQL DataFrame `df` with an array column named `scores` containing integers. They need to create a new column `passing` that is true only when every element in `scores` is greater than or equal to 70. Which Spark SQL higher-order function should they use?

Medium
276

A data scientist is using Spark SQL to compute a running total of sales amounts for each customer, ordered by transaction date. The query must return, for each row, the sum of all previous sales for that customer up to and including the current row. Which window specification should be used?

Medium
277

Which process is responsible for tracking the location of data blocks cached in the executors?

Medium
278

Which Spark component is responsible for maintaining the Directed Acyclic Graph (DAG) of stages and tasks?

Easy
279

In a Databricks Spark cluster, which component is primarily responsible for scheduling tasks and managing the distribution of computation across the worker nodes?

Medium
280

Which command is used to display the logical and physical execution plans for a given Spark SQL query?

Easy
281

What is the consequence of having 'wide dependencies' in a Spark job regarding the Spark Architecture?

Hard
282

Which TWO of the following are benefits of using Delta Lake over standard Parquet files in Spark SQL?

Medium
283

A data engineer runs a Spark job on a Databricks cluster using the default FIFO scheduler. They notice that a long-running job is holding all cluster resources, and short ad-hoc queries submitted later are stuck waiting. The engineer wants to allow concurrent scheduling of multiple jobs within the same Spark application so that short jobs can run while the long job is still executing. Which Spark configuration should be set to enable this behavior?

Medium
284

A developer is migrating a local pandas script to the Pandas API on Spark. The dataset is large and partitioned across many executors. The developer executes a custom row-wise operation using a standard Python lambda function inside a `.apply()` method without specifying return types or using vectorized operations. Why might this approach cause performance degradation in Databricks?

Medium
285

A developer is using Spark SQL to analyze a DataFrame that contains a column named tags, which holds an array of strings for each row. The developer needs to filter rows where the array contains the string 'urgent' and also produce a new column with the number of elements in the array. Which TWO Spark SQL expressions should be used in the query? (Choose two.)

Medium
286

In the context of the Spark Driver, what is the 'DAG' and why is it important?

Easy
287

A data engineer is building a Spark SQL pipeline that must return the top 3 highest-paid employees within each department from a Delta table named `employees` with columns `dept`, `name`, and `salary`. The engineer wants a single query that produces one row per qualifying employee, ranked by salary descending within each department, without collapsing rows. Which approach should be used?

Medium
288

A developer notices that a Spark DataFrame job on Databricks is running slowly and the Spark UI shows that many tasks are reading from a Delta table with a large number of small files. The job performs a filter on a date column and then aggregates results. Which optimization technique will most directly improve read performance in this scenario?

Easy
289

You are troubleshooting a Structured Streaming job that reads from Kafka and writes to a Delta table using foreachBatch. The job occasionally processes the same Kafka offsets twice after a task retry, causing duplicate rows in the Delta table. You want to ensure that each micro-batch's output is applied exactly once. Which change should you make inside the foreachBatch function?

Hard
290

What is the function of the 'Executor' within the Spark execution model?

Medium
291

You are running a Structured Streaming query on Databricks that reads from a Kafka topic and writes to a Delta table. The query uses `option("maxOffsetsPerTrigger", 10000)` to limit the number of records per micro-batch. During a peak, the Kafka topic accumulates a large backlog. You notice that the query is processing data but the backlog is not decreasing. What is the most likely cause?

Hard
292

Refer to the exhibit. Which of the following is the most likely cause for this 'shuffle fetch failure' in a Databricks cluster?

Medium
293

Which of the following describes the 'Driver' process in a Spark application?

Easy
294

A data engineer has a Pandas-on-Spark DataFrame `psdf` with a column `event_ts` stored as string timestamps. They run `psdf['event_ts'].astype('datetime64[ns]')` and then call `.dt.hour` on the resulting Series. In a Databricks notebook, what is the result of this operation?

Medium
295

You are debugging a PySpark DataFrame job on Databricks that performs multiple transformations and actions on a large delta table. You notice that the execution plan shows redundant computations where the same upstream DataFrame is evaluated repeatedly. Which transformation should you apply to optimize this workflow and avoid recomputing the upstream lineage?

Medium

Frequently asked questions

What does the scenario questions domain cover on the Databricks-Spark-Assoc exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 295 scenario questions questions in the Databricks-Spark-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.