Courseiva
← Back to Databricks Certified Associate Developer for Apache Spark questions

Scenario-based practice

Hard Difficulty Questions

Practise Databricks Certified Associate Developer for Apache Spark practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
Databricks-Spark-Assoc
exam code
Databricks
vendor

Scenario guide

How to approach hard difficulty questions

These are the questions most candidates get wrong. They require connecting multiple concepts, reading tricky output, or knowing edge-case behaviour that isn't on most study cards. Practising them trains you to operate under uncertainty — a necessary skill on the real exam.

Quick answer

Hard Difficulty Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related Databricks-Spark-Assoc topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?

Exhibit

24/05/20 10:00:00 INFO TaskSchedulerImpl: Adding task set 1.0 with 200 tasks
24/05/20 10:05:00 WARN TaskSetManager: Lost task 15.0 in stage 1.0: ExecutorLostFailure
Question 2hardmultiple choice
Full question →

You are building a Structured Streaming pipeline that reads from a Delta table source and applies a stateful deduplication using dropDuplicates on a composite key. After several hours, the job fails with an error indicating that the state store has grown too large. You need to bound the state size while still removing duplicate events that arrive within a reasonable window. Which approach should you take?

Question 3hardmultiple choice
Full question →

A Spark application reads a large Parquet dataset, performs a filter, and then writes the result to Delta Lake. The job fails with an OutOfMemoryError on the driver. The driver heap size is set to 4 GB, and the application uses a broadcast join for a small lookup table. Which action is most likely to resolve the driver OOM while preserving the broadcast join?

Question 4hardmulti select
Full question →

Which THREE of the following are valid ways to create a DataFrame from an existing table in Spark SQL?

Question 5hardmultiple choice
Full question →

A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes it locally. The developer wants to avoid the driver OOM while still obtaining the results. Which approach is most appropriate?

Question 6hardmulti select
Full question →

A Spark application is running on a Databricks cluster with adaptive query execution (AQE) enabled. The job reads a partitioned Parquet dataset, performs a join between two large tables, and then writes the result. The Spark UI shows that some tasks are taking significantly longer than others, and there is a high degree of skew in the join keys. Which TWO actions can help mitigate the skew and improve performance? (Choose two.)

Question 7hardmultiple choice
Full question →

A developer is using Structured Streaming with a Kafka source and wants to ensure that each message is processed exactly once, even in the event of failures. The developer has set a checkpoint location and is using `foreachBatch` to write to an external database. Which additional step is necessary to achieve exactly-once semantics?

Question 8hardmultiple choice
Full question →

A data engineer is using Spark SQL to join two large tables, 'orders' and 'customers', on customer_id. The 'orders' table has a column 'order_date' and the 'customers' table has a column 'signup_date'. The engineer wants to include only orders placed after the customer's signup date. Which join condition correctly implements this requirement?

Question 9hardmultiple choice
Full question →

A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?

Question 10hardmulti select
Full question →

A developer is building a Spark Connect client application that will run against a remote Databricks cluster. The application must handle data locally for small results and must avoid unsupported client APIs. Which two statements correctly describe behavior or limitations the developer should account for? (Choose two.)

Question 11hardmulti select
Full question →

A data engineer is using Spark Connect to interact with a remote Databricks cluster. The engineer needs to understand the limitations of Spark Connect compared to traditional Spark. Which TWO of the following statements accurately describe limitations of Spark Connect? (Choose two.)

A developer is using Spark Connect to run a complex job that includes a Python UDF. The job fails with a serialization error. The developer suspects that the UDF is trying to use a client-side library that is not available on the cluster. Which action should the developer take to resolve this?

Question 13hardmulti select
Full question →

A data scientist is working with a Pandas API on Spark DataFrame `psdf` that contains a column `timestamp` of type `TimestampType`. They need to extract the day of the week and the hour of the day as new integer columns. Which two methods correctly achieve this while preserving distributed execution? (Choose two.)

Question 14hardmultiple choice
Full question →

A data scientist is using Pandas API on Spark to analyze a large dataset. They perform a groupby operation followed by an apply of a custom function that returns a pandas Series. The operation is running very slowly and sometimes fails with out-of-memory errors on the executors. What is the most likely cause and the recommended solution?

Question 15hardmultiple choice
Full question →

A data engineer has a DataFrame `events` with columns `user_id` and `ts`. They need to add a column `prev_ts` that holds the previous event timestamp for each user, ordered by `ts` ascending, without collapsing rows. Which operation accomplishes this?

Question 16hardmulti select
Full question →

A Databricks job processing a Delta table with 50,000 small files is slow because each task opens many files. The developer wants to reduce the number of files read per task without rewriting the table. Which two actions should be taken? (Choose two.)

Question 17hardmulti select
Full question →

Which TWO of the following statements accurately describe the role of the Spark Executor in a cluster deployment?

Question 18hardmultiple choice
Full question →

A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?

Question 19hardmultiple choice
Full question →

A streaming job reads clickstream events and performs a stateful `groupBy("sessionId").count()` with a 15-minute watermark on `eventTime`. The developer notices that state store size keeps growing without bound even though watermarks are configured. Which change most directly limits the state growth?

Question 20hardmulti select
Full question →

A developer must remove duplicate rows from a DataFrame `orders` so that only the most recent order per `customer_id` is kept, based on `order_ts`. Which TWO approaches achieve this correctly? (Choose two.)

These Databricks-Spark-Assoc practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-Spark-Assoc questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.