Courseiva
← Back to Databricks Certified Associate Developer for Apache Spark questions

Scenario-based practice

Select Two (Multi-Select) Questions

Practise Databricks Certified Associate Developer for Apache Spark practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
Databricks-Spark-Assoc
exam code
Databricks
vendor

Scenario guide

How to approach select two (multi-select) questions

Multi-select questions tell you to 'Choose TWO' or 'Choose THREE'. Getting partial credit is not a thing — you must select all correct answers with no incorrect ones. The stem always states how many to choose, so trust it. These questions require precision, not best-guess elimination.

Quick answer

Select Two (Multi-Select) Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related Databricks-Spark-Assoc topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmulti select
Full question →

Which THREE of the following are valid ways to create a DataFrame from an existing table in Spark SQL?

Question 2mediummulti select
Full question →

Which THREE of the following sources support streaming read operations in Spark Structured Streaming?

A data engineer is using Spark Connect from a remote Python client to interact with a Databricks cluster. The engineer wants to understand which operations are executed on the server side versus the client side. Which two statements correctly describe this behavior? (Choose two.)

Question 4mediummulti select
Full question →

A developer is working with a Pandas API on Spark DataFrame `psdf` and wants to perform operations that are efficient in a distributed environment. Which two operations are considered efficient and do not require collecting data to the driver? (Choose two.)

Question 5hardmulti select
Full question →

Which TWO of the following statements accurately describe the role of the Spark Executor in a cluster deployment?

Question 6mediummulti select
Full question →

A developer is using Spark SQL to analyze a DataFrame that contains a column named tags, which holds an array of strings for each row. The developer needs to filter rows where the array contains the string 'urgent' and also produce a new column with the number of elements in the array. Which TWO Spark SQL expressions should be used in the query? (Choose two.)

Question 7hardmulti select
Full question →

A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?

Question 8hardmulti select
Full question →

Which TWO of the following statements accurately describe the relationship between Spark Executors and memory management within a Databricks cluster?

Question 9mediummulti select
Full question →

Which TWO factors influence the effective parallelism of a Spark application?

Question 10hardmulti select
Full question →

Which THREE of the following are valid ways to monitor or debug Spark SQL query performance in Databricks?

Question 11hardmulti select
Full question →

A Databricks engineer is diagnosing why a Spark job's shuffle phase writes a very large amount of data to disk. The engineer wants to reduce shuffle overhead by changing how the job is structured and configured. Which TWO actions are most likely to reduce the volume of shuffle data written? (Choose two.)

Question 12mediummulti select
Full question →

A Spark application is running on a cluster with 5 executors. The driver program creates a broadcast variable that is used in a transformation. Which two components are directly involved in distributing and using the broadcast variable? (Choose two.)

Question 13mediummulti select
Full question →

A data engineer has a PySpark DataFrame `events` with columns `user_id`, `event_time` (timestamp), and `payload` (string). They must produce a new DataFrame where each row is enriched with the `payload` value from the user's immediately preceding event, ordered by `event_time`, without collapsing any rows. Which TWO approaches accomplish this? (Choose two.)

Question 14mediummulti select
Full question →

A developer is using Spark on Databricks and notices that a particular job has many stages due to shuffle operations. They want to understand the role of the shuffle in the Spark execution model. Which two statements accurately describe the behavior of a shuffle operation in Spark? (Choose two.)

Question 15mediummulti select
Full question →

A data engineer is writing a Spark SQL query that joins a `transactions` table to a `customers` table on `customer_id`. The engineer wants to ensure that rows from `transactions` with no matching customer are still returned, with nulls for customer columns, and also wants to exclude duplicate rows that arise from the join. Which TWO clauses should the engineer include? (Choose two.)

Question 16mediummulti select
Full question →

A developer is troubleshooting a Spark Connect client that intermittently fails with connection errors to a Databricks cluster. Which two configuration practices help ensure stable connectivity? (Choose two.)

Question 17mediummulti select
Full question →

A developer writes a Spark SQL query that groups orders by region and computes the total revenue per region, but also needs to return the number of distinct customers per region in the same result set. Which TWO expressions correctly compute the distinct customer count per region in a single GROUP BY region query? (Choose two.)

Question 18hardmulti select
Full question →

You are using Spark SQL to join two large Delta tables, orders and customers, on a common column customer_id. The orders table is partitioned by order_date, and the customers table is not partitioned. You need to ensure the join is efficient and minimizes shuffling. Which TWO actions should you take? (Choose two.)

Question 19mediummulti select
Full question →

A developer is building a Structured Streaming job that reads from a Delta table as a stream and writes to another Delta table. The job must support exactly-once processing and allow the output to be updated incrementally. Which two options are required to achieve exactly-once semantics? (Choose two.)

Question 20mediummulti select
Full question →

A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)

These Databricks-Spark-Assoc practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-Spark-Assoc questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.