Be able to read Databricks error messages, inspect job run details and Spark UI, and apply the correct fix: grant Unity Catalog privileges, configure cloud storage access, or address skew with salting, AQE, or repartitioning. The most important thing is matching the symptom to the exact Databricks feature or permission.
Start practicing
Debugging and Deploying — choose a session length
Free · No account required
Domain overview
This domain covers debugging and deploying Databricks workloads: diagnosing job failures across environments, resolving Unity Catalog permission errors, fixing cloud storage access issues, and tuning skewed queries. Questions present realistic error messages, exhibits, or symptoms and ask you to select the correct Databricks feature, configuration, or fix.
Exam objectives
Use Databricks job run history, Spark UI, and cluster logs to diagnose failures.
Compare environments using Databricks workspace, cluster, and job configuration differences.
Resolve Unity Catalog PERMISSION_DENIED errors with GRANT USE CATALOG/SCHEMA/TABLE.
Fix data skew using salting, AQE skew join, or repartitioning before writes.
Assuming a production failure is code-related when it is actually a missing Unity Catalog grant or cloud IAM role difference.
Ignoring that AQE skew join handles join skew but not skew during groupBy or write operations.
Confusing storage credential and external location permissions with Unity Catalog table grants when debugging S3 403 errors.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data engineer is troubleshooting a Databricks Workflow where a downstream task relies on an upstream task's output. Which TWO actions ensure the data dependency is correctly handled during a failure scenario?
2Refer to the exhibit. A data engineer is deploying a production pipeline that references a table in the default schema. The job fails with the provided error. What is the root cause?
3A data engineer is automating the deployment of Databricks assets using CI/CD. The pipeline fails because the 'databricks-cli' command cannot find the workspace. What is the most likely cause?
4You are monitoring a long-running Databricks job. You notice that the memory usage on the driver node is steadily increasing until it crashes. Which debugging action is most appropriate?
5Which THREE strategies should a data engineer use to optimize the debugging of failed production Databricks Jobs?
6Where can a data engineer find the standard output and error logs for a specific task within a Databricks Workflow?
7A data engineer is debugging a slow-running query. They notice that the data is skewed, causing one task to take significantly longer than others. Which approach effectively addresses this skew?
8Refer to the exhibit. A Databricks job fails with a 403 Forbidden error when trying to write to the S3 bucket. Why does this happen?
9A production Databricks workflow involves a task that runs a notebook. The notebook takes 15 minutes to finish, but the workflow is set to timeout after 10 minutes. What happens?
10A data engineer is investigating a job failure that occurred only in the production environment. Which TWO features in Databricks help in comparing the production environment to the development environment?
11Refer to the exhibit. Why is Task D marked as 'Skipped'?
12What is the primary benefit of using a Job Cluster instead of an All-Purpose Cluster for production workloads?
13Refer to the exhibit. Which action is the most appropriate to resolve this memory-related failure during the job execution?
14A data engineer is preparing to deploy a production Databricks workflow. Which TWO best practices should be implemented to ensure maintainability and robust error handling?
15Refer to the exhibit. Which configuration change is required to enable multiple instances of this job to run simultaneously?
16A data engineer needs to inspect the logs of a long-running Databricks job that has already completed. Where should they navigate in the Databricks UI to find the driver logs?
17A data engineer is deploying a Databricks job using Databricks Asset Bundles. They want to ensure that the job uses a specific cluster configuration that is defined once and reused across multiple tasks. Which bundle feature should they use?
18A data engineer is using Databricks Asset Bundles to deploy a job that runs a Python wheel task. The bundle is deployed to a production workspace using a service principal. The job fails with the error: `Library installation failed for library due to user error: Could not find wheel file`. The engineer confirms the wheel file exists in the bundle's `dist` folder. What is the most likely cause of this failure?
19A data engineer is debugging a Databricks job that fails with a `SparkException: Job aborted due to stage failure` in production. They need to identify the root cause. Which two actions should they take to gather relevant diagnostic information? (Choose two.)
20A data engineer is deploying a Databricks job that uses a Python wheel task. The job fails with the error: 'ModuleNotFoundError: No module named 'my_library''. The wheel file is stored in DBFS at 'dbfs:/FileStore/wheels/my_library-0.1.0-py3-none-any.whl'. The job cluster is configured with a cluster policy that restricts library installation from DBFS. What is the most likely cause of the failure?
21A data engineer deploys a Databricks Job that runs a notebook task. The notebook writes to a Delta table in Unity Catalog. The job fails with the error: 'PERMISSION_DENIED: User does not have USE CATALOG on catalog 'prod'.' The engineer confirms the job's service principal has USE CATALOG granted on the catalog. Which configuration should the engineer check next?
22A data engineer is troubleshooting a Databricks job that fails with a `SparkException: Job aborted due to stage failure` and the error log shows `java.lang.OutOfMemoryError: GC overhead limit exceeded` on an executor. The job processes a large dataset using a `groupByKey` operation. Which action should the engineer take to resolve the issue while minimizing changes to the existing code?
23A data engineer is using Databricks Repos to manage code for a production job. They need to ensure that the job always uses the latest version of the code from a specific branch. Which Git operation should they perform before running the job?
24A data engineer needs to grant a group of users the ability to run a specific Databricks job but not modify its configuration. The job is managed by a service principal. Which permission level should be assigned to the group on the job?
25A data engineer is deploying a Databricks Asset Bundle (DAB) that defines a job with a notebook task. The bundle validates locally, but deployment fails with 'Error: cannot find notebook at path /Workspace/Users/dev@example.com/pipeline/ingest'. The engineer confirms the notebook exists in the workspace at that exact path. Which action should the engineer take to resolve the deployment failure?
26A data engineer needs to ensure that a Databricks job can be retried automatically if it fails due to a transient cluster error. The job is configured with a maximum of 3 retries. After the third failure, the engineer wants to receive an email notification. Which feature should be used to accomplish this?
27A data engineer is using Databricks Repos to manage a project. They need to ensure that the production job always uses the code from the 'main' branch, even if developers push changes to other branches. Which Git reference should be used in the job configuration to achieve this?
28A data engineer is troubleshooting a Databricks job that fails with a 'TaskFailed' error. The job uses a cluster with autoscaling enabled. The engineer suspects that the failure is due to memory issues on the workers. Which TWO actions should the engineer take to diagnose and resolve the issue? (Choose two.)
29A data engineer is debugging a Databricks job that reads from a Delta table and writes to another Delta table. The job occasionally fails with 'ConcurrentAppendException'. The engineer wants to minimize failures while maintaining data correctness. Which approach should the engineer take?
30A data engineer is preparing to deploy a production Databricks Workflow that must be maintainable and auditable. The engineer wants to ensure that changes to the workflow are tracked and that failures can be diagnosed quickly. Which TWO practices should the engineer implement? (Choose two.)
31A data engineer is configuring a Databricks Workflow that must run a notebook task only after a previous task that writes to a Delta table has completed successfully. The engineer wants to ensure that if the first task fails, the second task does not run. Which feature should the engineer use to define this dependency?
Be able to read Databricks error messages, inspect job run details and Spark UI, and apply the correct fix: grant Unity Catalog privileges, configure cloud storage access, or address skew with salting, AQE, or repartitioning. The most important thing is matching the symptom to the exact Databricks feature or permission.
The Courseiva Databricks-DE-Pro question bank contains 31 questions in the Debugging and Deploying domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Debugging and Deploying domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included