Courseiva
← Back to Databricks Certified Associate Developer for Apache Spark questions

Scenario-based practice

Refer to the Exhibit Practice Questions

Practise Databricks Certified Associate Developer for Apache Spark practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

10
scenario questions
Databricks-Spark-Assoc
exam code
Databricks
vendor

Scenario guide

How to approach refer to the exhibit practice questions

Practise exhibit-style questions that ask you to read a topology, table, command output or diagram before choosing the best answer.

Quick answer

Exhibit-style questions test whether you can read a topology, command output, diagram or table before choosing the best answer.

How to extract the relevant detail from an exhibit.

How topology, command output or routing information affects the answer.

How to avoid answering from memory before reading the evidence.

How to map the exhibit back to the exam objective.

Related practice questions

Related Databricks-Spark-Assoc topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?

Exhibit

24/05/20 10:00:00 INFO TaskSchedulerImpl: Adding task set 1.0 with 200 tasks
24/05/20 10:05:00 WARN TaskSetManager: Lost task 15.0 in stage 1.0: ExecutorLostFailure
Question 2mediummultiple choice
Study the full Python automation breakdown →

Refer to the exhibit.

Traceback (most recent call last): File "app.py", line 12, in <module> df = spark.read.table("default.sales") File "/opt/spark/python/pyspark/sql/session.py", line 314, in table

return DataFrame(self._client.execute_plan(parser.parse_table(name))))

File "/opt/spark/python/pyspark/sql/connect/client/core.py", line 112, in execute_plan(y+"sessionID"), grpc.RpcError: StatusCode.UNAVAILABLE

An engineer attempts to run a PySpark script using Spark Connect but encounters the traceback shown above. What is the most likely root cause of this execution failure?

Question 3hardmultiple choice
Full question →

Refer to the exhibit. What is the best way to resolve this error?

Exhibit

Error: org.apache.spark.SparkException: Job aborted due to stage failure: Total size of serialized results of 5000 tasks (2048 MB) is bigger than spark.driver.maxResultSize (1024 MB)
Question 4mediummultiple choice
Full question →

Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?

Exhibit

24/05/10 10:00:00 WARN DAGScheduler: Broadcasting large task binary with size 12.5 MiB
Question 5mediummultiple choice
Full question →

Refer to the exhibit. Which of the following is the most likely cause for this 'shuffle fetch failure' in a Databricks cluster?

Exhibit

org.apache.spark.SparkException: Job aborted due to stage failure: Task ... in stage 1.0 (TID 1) had a shuffle fetch failure
Question 6mediummultiple choice
Full question →

Based on the exhibit, what is the most likely reason for this error in a Databricks notebook?

Exhibit

Refer to the exhibit.

[ERROR LOG]
org.apache.spark.SparkException: Job aborted due to stage failure: 
Task 0 in stage 1.0 failed 4 times, last failure: 
ResultTask(0, 4) failed: 
java.lang.UnsupportedOperationException: 'to_pandas' is not supported on a DataFrame with a very large size.
[END OF EXHIBIT]
Question 7hardmultiple choice
Full question →

Refer to the exhibit. Which configuration issue is most likely causing this warning in a Spark application?

Exhibit

24/05/10 10:00:00 WARN TaskSchedulerImpl: Initial job has not accepted any resources; check your cluster UI to ensure that workers are registered and have sufficient resources.
Question 8hardmultiple choice
Full question →

Refer to the exhibit. Based on the error log, what is the most likely cause of the job failure?

Exhibit

org.apache.spark.SparkException: Job aborted due to stage failure: Task 5 in stage 1.0 failed 4 times
... executor lost: ExecutorLostFailure (executor 2 exited caused by one of the running tasks) Reason: Container killed by YARN for exceeding memory limits.
Question 9mediummultiple choice
Full question →

Refer to the exhibit. Which action is the most likely cause of the error shown in the Spark job logs?

Exhibit

Error: java.lang.OutOfMemoryError: Java heap space
State: FAILED
Task: 452
Stage: 12
Action: collect()
Question 10hardmultiple choice
Full question →

Refer to the exhibit. Which performance indicator suggests that Task 15 is likely causing a performance bottleneck during the execution of a join operation?

Exhibit

Task 12: 500MB (Local Disk)
Task 15: 1.2GB (Shuffle Read)
Task 22: 400MB (Shuffle Write)

These Databricks-Spark-Assoc practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-Spark-Assoc questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.