Sample questions
Databricks Certified Associate Developer for Apache Spark practice questions
Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?
You are building a Structured Streaming pipeline that reads from a Delta table source and applies a stateful deduplication using dropDuplicates on a composite key. After several ho…
What happens when an action is called on a Spark DataFrame?
A data engineer is working with the Pandas API on Spark and needs to convert a Spark DataFrame named `sdf` into a pandas DataFrame so it can be processed locally on the driver node…
Which clause is used in a SELECT statement to filter the results based on aggregated values?
You are processing a streaming dataset of sensor readings. You need to calculate the average temperature every 10 minutes, allowing data to arrive up to 2 minutes late. Which windo…
Which THREE of the following are valid ways to create a DataFrame from an existing table in Spark SQL?
A data engineer submits a Spark application using spark-submit in client deploy mode from an edge node. The application reads a large Parquet dataset, performs a groupBy aggregatio…
Which component in the Spark architecture is responsible for maintaining the state of the Spark application and coordinating the execution of tasks across the cluster?
You are processing a large dataset in Spark SQL and need to ensure that small files are avoided when writing data to Delta Lake. Which approach effectively minimizes small file gen…
Which of the following Spark SQL configuration settings should be adjusted to prevent the 'Driver OOM' error when collecting massive amounts of query results to the driver node?
Which of the following describes the behavior of a 'Broadcast Hash Join' in Spark SQL?
A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes i…
Which THREE of the following sources support streaming read operations in Spark Structured Streaming?
What is the result of applying the COALESCE function in Spark SQL when multiple arguments are provided?
A data engineer has a Pandas-on-Spark DataFrame `psdf` with a column `event_time` stored as string. They run `psdf['event_time'] = pd.to_datetime(psdf['event_time'])` where `pd` is…
A developer is using Structured Streaming with a Kafka source and wants to ensure that each message is processed exactly once, even in the event of failures. The developer has set…
Refer to the exhibit. Traceback (most recent call last): File "app.py", line 12, in <module> df = spark.read.table("default.sales") File "/opt/spark/python/pyspark/sql/ses…
When working with Delta Lake tables in Databricks, which command should you use to optimize the physical layout of files to improve query performance?
A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch dur…
A data engineer is using Spark Connect from a remote Python client to interact with a Databricks cluster. The engineer wants to understand which operations are executed on the serv…
You are writing a Databricks notebook and want to use the Pandas API on Spark. Which import statement should you use to access the Pandas API on Spark?
A developer writes a Structured Streaming query that reads from a Kafka topic with `spark.readStream.format("kafka")` and then calls `.writeStream.format("console").start()`. The q…
Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize…
Troubleshooting and Tuning DataFrame AppsmediumSee the answer and why each option is right or wrong →