Courseiva

CCNA Debugging and Deploying Questions

31 questions · Debugging and Deploying · All types, answers revealed

1
MCQhard

Refer to the exhibit. Which action is the most appropriate to resolve this memory-related failure during the job execution?

A.Decrease the number of partitions using spark.sql.shuffle.partitions.
B.Upgrade the cluster to an instance type with higher memory capacity.
C.Enable autoscaling to add more nodes to the cluster.
D.Change the file format from Parquet to CSV for better performance.
AnswerB

Upgrading to a memory-optimized instance family directly addresses the physical memory constraint identified in the logs. By increasing the memory allocated to each container, the executors can accommodate the memory footprint required for the transformation, preventing the container from being killed by the underlying resource manager.

Why this answer

The error indicates an Out of Memory (OOM) condition where the container exceeded its allocated memory limit. Increasing the instance type to a memory-optimized VM size provides more RAM per executor, allowing Spark to process larger partitions without triggering the YARN memory killer. This is a common bottleneck in memory-intensive operations like wide transformations, where adjusting the cluster configuration is more effective than attempting to optimize complex query logic.

Exam trap

Candidates frequently try to solve memory errors by tuning Spark shuffle partitions or changing code logic, ignoring that an outright hardware memory ceiling requires a memory-optimized instance type.

2
MCQeasy

A data engineer needs to ensure that a Databricks job can be retried automatically if it fails due to a transient cluster error. The job is configured with a maximum of 3 retries. After the third failure, the engineer wants to receive an email notification. Which feature should be used to accomplish this?

A.Set up a Databricks SQL alert on the job's run history table.
B.Configure the job to use a job cluster with autoscaling and set a retry policy.
C.Use a webhook to call an external service that sends an email.
D.Configure the job's `email_notifications` with `on_failure` and set the email address.
AnswerD

The `email_notifications` setting in a Databricks job allows you to specify email addresses to notify on failure. When combined with retries, the notification is sent after all retries are exhausted (i.e., after the final failure). This is the standard way to alert on job failure. It requires no additional infrastructure and is built into the job configuration.

Why this answer

Databricks jobs support built-in email notifications that can be triggered on failure. By configuring `email_notifications` with `on_failure`, the specified recipients receive an email after the job fails, including after all retries are exhausted. This is the most direct and native method.

Other options require external services or are not designed for this purpose.

Exam trap

The trap here is overlooking the native email notification feature and instead considering external alerting mechanisms that add unnecessary complexity.

3
MCQmedium

You are monitoring a long-running Databricks job. You notice that the memory usage on the driver node is steadily increasing until it crashes. Which debugging action is most appropriate?

A.Increase the number of worker nodes in the cluster.
B.Use the Spark UI to identify tasks using 'collect()' or 'toPandas()' on large datasets.
C.Update the cluster's Spark configuration to disable the driver's log monitoring.
D.Lower the 'spark.driver.maxResultSize' setting in the cluster configuration.
AnswerB

The Spark UI allows you to inspect the execution plan and identify operations that move data from worker nodes to the driver node. Using 'collect()' on large datasets is a classic cause of driver out-of-memory errors. Identifying these bottlenecks allows you to replace them with distributed data writing.

Why this answer

Driver memory crashes are typically caused by collecting large datasets from workers to the driver. By identifying the 'collect' or 'toPandas' operations, you can optimize the code to process data in a distributed manner. This is critical for scaling data pipelines; understanding why the driver is overloaded ensures that jobs remain performant as data volumes grow and prevents common failures in production batch processing environments.

Exam trap

Candidates often suggest increasing the driver instance size as the primary solution. This ignores the root cause, which is an architectural flaw in the code pulling data into memory.

4
MCQmedium

Refer to the exhibit. Why is Task D marked as 'Skipped'?

A.Task D reached its execution timeout limit before it could start.
B.Task D has a dependency on Task B, which failed.
C.The cluster running Task D was terminated by an administrator.
D.Task D was manually disabled in the job configuration.
AnswerB

The workflow engine marks tasks as 'Skipped' if they depend on an upstream task that failed. Since Task B failed, Task D, which relies on the successful completion of Task B, is automatically skipped to prevent the execution of a potentially inconsistent or broken data process.

Why this answer

In Databricks Workflows, a task is 'Skipped' if its upstream dependencies have not been met, typically because a required parent task failed. Because Task B failed, the downstream Task D—which likely depends on Task B—cannot execute. This mechanism prevents the workflow from processing invalid or incomplete data, maintaining the integrity of the data pipeline and preventing downstream failures from compounding existing upstream errors.

Exam trap

Candidates mistakenly believe a 'Skipped' task status means the task was manually disabled, confusing it with upstream dependency failures where parent task failure cascades downstream.

5
MCQeasy

A data engineer needs to inspect the logs of a long-running Databricks job that has already completed. Where should they navigate in the Databricks UI to find the driver logs?

A.The 'Clusters' tab under the 'Event Log' section.
B.The 'Runs' tab within the specific Job's detail page.
C.The 'Workspace' folder where the notebook is stored.
D.The 'SQL Warehouses' monitoring dashboard.
AnswerB

The 'Runs' tab contains the history of all past job executions. By clicking into a specific run, users can access the task logs, driver logs, and Spark UI links. This is the intended UI location for inspecting logs for completed or failed jobs to determine the root cause.

Why this answer

The Job Runs view provides a comprehensive audit trail of all executions, including specific task logs and driver output. Accessing the 'Runs' tab for a specific job allows the engineer to select a historical run and view the standard output and standard error streams. This is the primary interface for post-mortem debugging and understanding the execution context of failed tasks that are no longer actively running on a cluster.

Exam trap

Candidates often look in the 'Cluster' logs tab or the 'Workspace' file browser, assuming job logs are stored in the same place as cluster-level event logs or source code.

6
MCQeasy

A data engineer is deploying a Databricks job using Databricks Asset Bundles. They want to ensure that the job uses a specific cluster configuration that is defined once and reused across multiple tasks. Which bundle feature should they use?

A.Use a `new_cluster` definition for each task, ensuring they all have identical configurations.
B.Define a `job_cluster` in the job's `job_clusters` section and reference it by key in each task.
C.Define a cluster in the `resources/clusters` directory and reference it in the job.
D.Use an `existing_cluster_id` for all tasks, pointing to a pre-created cluster.
AnswerB

The `job_clusters` section allows defining reusable cluster configurations at the job level. Each task can then reference a cluster by its key using the `job_cluster_key` field. This promotes consistency and reduces duplication. It is the intended way to share cluster settings across tasks within a single job.

Why this answer

The `job_clusters` section in a Databricks job definition allows specifying cluster configurations that can be referenced by multiple tasks via `job_cluster_key`. This is the correct way to define a cluster once and reuse it across tasks within the same job, ensuring consistency and simplifying maintenance. Other options either duplicate configuration or rely on external resources.

Exam trap

The trap here is confusing job-level cluster definitions with task-level `new_cluster` or external cluster references, which do not provide the same reusability.

7
MCQmedium

A data engineer is using Databricks Repos to manage a project. They need to ensure that the production job always uses the code from the 'main' branch, even if developers push changes to other branches. Which Git reference should be used in the job configuration to achieve this?

A.The branch name 'main'
B.A specific commit hash
C.The remote URL of the repository
D.A tag named 'production'
AnswerA

Specifying the branch name 'main' in the job configuration ensures that the job checks out the latest commit on that branch at runtime. This means any new commits pushed to 'main' will be used in subsequent job runs, satisfying the requirement to always use production code from 'main'.

Why this answer

In Databricks Repos, job configurations can reference a Git branch, tag, or commit. To always use the latest code from the 'main' branch, the job should be configured with the branch name 'main'. At runtime, Databricks checks out the latest commit on that branch.

Using a commit hash or tag would pin to a specific version and not automatically update, which is unsuitable for continuous deployment.

Exam trap

The trap here is assuming that a tag or commit hash provides the same automatic updates as a branch, when in fact they are static references.

8
MCQeasy

A data engineer is using Databricks Repos to manage code for a production job. They need to ensure that the job always uses the latest version of the code from a specific branch. Which Git operation should they perform before running the job?

A.Merge the remote branch into the local branch using the Repos UI.
B.Pull the latest changes from the remote repository into the Repo.
C.Create a new branch from the current branch and switch to it.
D.Commit and push local changes to the remote repository.
AnswerB

Pulling updates the local Repo with the latest commits from the remote branch. This ensures the job runs the most recent code. In Databricks Repos, the pull operation is available via the UI or API and is a standard step before running jobs that depend on the latest code.

Why this answer

To ensure a Databricks job uses the latest code from a Git branch, the engineer must pull the latest changes into the Repo. This updates the working directory with the most recent commits. Committing, branching, or merging are not appropriate for simply retrieving updates.

Pulling is the standard Git operation for this purpose.

Exam trap

The trap here is confusing push with pull; pushing sends local changes to remote, while pulling retrieves remote changes to local.

9
Multi-Selecthard

A data engineer is troubleshooting a Databricks Workflow where a downstream task relies on an upstream task's output. Which TWO actions ensure the data dependency is correctly handled during a failure scenario?

Select 2 answers
A.Configure the downstream task to use the 'depends_on' attribute to reference the upstream task ID.
B.Set the task timeout to zero to prevent the workflow from ever stopping during a failure.
C.Enable the 'repair and rerun' feature to target only the failed tasks in the pipeline.
D.Hardcode the file path of the upstream output into the downstream task configuration.
E.Disable all retries to ensure that the error log is captured immediately upon failure.
AnswersA, C

The 'depends_on' attribute explicitly defines the task DAG structure within Databricks Workflows. By creating this dependency, the scheduler guarantees that the downstream task will only execute if the upstream task succeeds, preventing erroneous runs when input data is missing or corrupted due to preceding failures.

Why this answer

Proper dependency management ensures that downstream tasks do not attempt to process incomplete or missing data from upstream failures. By using task dependencies and repair-and-rerun capabilities, engineers can isolate failures and maintain data integrity. These features are critical in production pipelines to prevent downstream corruption and ensure that data lineage remains consistent throughout the workflow execution cycle regardless of individual component failures.

Exam trap

Candidates often confuse workflow task configuration attributes like 'depends_on' with general cluster settings, or forget that 'repair and rerun' requires specific task targeting instead of restarting the entire pipeline from scratch.

10
MCQhard

A data engineer is troubleshooting a Databricks job that fails with a `SparkException: Job aborted due to stage failure` and the error log shows `java.lang.OutOfMemoryError: GC overhead limit exceeded` on an executor. The job processes a large dataset using a `groupByKey` operation. Which action should the engineer take to resolve the issue while minimizing changes to the existing code?

A.Increase the executor memory by adding `spark.executor.memory=16g` to the cluster Spark configuration.
B.Set `spark.sql.shuffle.partitions` to a higher value to increase the number of partitions after shuffle.
C.Replace `groupByKey` with `reduceByKey` to reduce memory pressure by combining values locally before shuffling.
D.Switch the job to use a larger cluster with more worker nodes to distribute the load.
AnswerC

`groupByKey` shuffles all values for a key without local aggregation, which can cause excessive memory usage and GC overhead. `reduceByKey` performs map-side combining, reducing the amount of data shuffled and lowering memory pressure. This change directly addresses the root cause while requiring minimal code modification, as both are transformations on key-value pairs. It is the most effective fix for the described symptom.

Why this answer

The `groupByKey` operation shuffles all values for each key without local aggregation, which can lead to excessive memory consumption and GC overhead. Replacing it with `reduceByKey` enables map-side combining, significantly reducing the shuffle size and memory footprint. This is the most direct and code-minimal fix.

Other options either mask the symptom or do not address the root cause.

Exam trap

The trap here is thinking that adding more memory or nodes will solve an OOM caused by an inefficient shuffle operation like `groupByKey`.

11
MCQhard

A data engineer is deploying a Databricks job that uses a Python wheel task. The job fails with the error: 'ModuleNotFoundError: No module named 'my_library''. The wheel file is stored in DBFS at 'dbfs:/FileStore/wheels/my_library-0.1.0-py3-none-any.whl'. The job cluster is configured with a cluster policy that restricts library installation from DBFS. What is the most likely cause of the failure?

A.The job cluster does not have the 'pip' package manager installed.
B.The wheel file path is incorrect; it should be 'dbfs:/FileStore/wheels/my_library-0.1.0-py3-none-any.whl' without the 'dbfs:/' prefix.
C.The cluster policy prevents installing libraries from DBFS, so the wheel was not installed.
D.The wheel file is not compatible with the cluster's Python version.
AnswerC

Cluster policies can restrict library sources. If the policy disallows DBFS libraries, the wheel will not be installed, leading to the ModuleNotFoundError. This is the most likely cause because the error explicitly states the module is missing, and the policy restriction directly prevents installation. The engineer should either adjust the policy to allow DBFS or move the wheel to a supported location like Unity Catalog volumes or cloud storage.

Why this answer

The error indicates that the Python module 'my_library' is not available in the job cluster's environment. Since the wheel is stored in DBFS and the cluster policy restricts library installation from DBFS, the wheel was not installed. The engineer must either modify the cluster policy to allow DBFS libraries or relocate the wheel to a permitted source, such as a Unity Catalog volume or cloud storage, and then install it.

Exam trap

The trap here is focusing on the wheel's compatibility or path syntax instead of the cluster policy restriction that blocks installation from DBFS.

12
Multi-Selecthard

Which THREE strategies should a data engineer use to optimize the debugging of failed production Databricks Jobs?

Select 3 answers
A.Implement structured logging within the application code to track variable states.
B.Grant all developers full cluster permissions to access logs directly on the nodes.
C.Modularize code into libraries (wheels) to enable easier unit testing and local debugging.
D.Configure alerts on job failures to send notifications to a team Slack or email channel.
E.Run all production jobs as the root user to avoid permission-related errors.
AnswersA, C, D

Structured logging provides granular insight into the execution path and variable values, which are otherwise unavailable once a job fails. By writing these logs to a persistent sink, engineers can reconstruct the state of the application at the exact moment of failure, significantly reducing the time required for investigation.

Why this answer

Efficient debugging relies on observability, modular design, and access control. By leveraging logs, modularizing code, and ensuring proper access, engineers can quickly isolate root causes. These strategies are essential for minimizing Mean Time to Recovery (MTTR) in production.

Proactive monitoring and well-structured codebases allow for faster identification of failures, ensuring that business-critical pipelines remain operational and that the team can respond to incidents with precision and speed.

Exam trap

Candidates often suggest manually inspecting cloud provider logs (like CloudWatch) first. They overlook that Databricks Jobs UI and Git integration provide the specific context needed for application-level failures.

13
MCQeasy

Where can a data engineer find the standard output and error logs for a specific task within a Databricks Workflow?

A.In the Databricks Filesystem (DBFS) root directory under /logs.
B.By clicking the 'Logs' tab within the specific task run details in the Jobs UI.
C.By querying the 'sys.logs' table in the Unity Catalog.
D.In the cluster configuration's 'Advanced Options' tab.
AnswerB

The Jobs UI provides a 'Logs' tab that captures stdout, stderr, and log4j outputs for every task execution. This is the official, supported way to view execution logs within the Databricks workspace, allowing engineers to quickly debug failures without leaving the browser or using complex CLI commands.

Why this answer

Databricks provides a comprehensive UI that aggregates logs for every task execution. By navigating to the job run and selecting the specific task, the engineer can view stdout and stderr logs directly. This is the primary method for diagnosing runtime exceptions, syntax errors, or logic failures in automated data pipelines, facilitating rapid troubleshooting without needing external tools or direct node access.

Exam trap

Candidates often look for logs in the cluster driver tab or external cloud storage buckets. They miss that the Jobs UI provides a direct, aggregated view of task-specific logs.

14
MCQhard

Refer to the exhibit. Which configuration change is required to enable multiple instances of this job to run simultaneously?

A.Increase the timeout threshold for the job.
B.Set max_concurrent_runs to a value greater than 1 in the job settings.
C.Add more workers to the cluster assigned to the job.
D.Enable 'Retries' in the job configuration.
AnswerB

The max_concurrent_runs parameter controls how many instances of a job can execute at the same time. Setting this to a value higher than 1 allows the job to be triggered even if a previous run is still in progress, enabling parallel processing of independent data workflows.

Why this answer

The error message explicitly states that the job does not allow concurrent runs. By default, many jobs are configured for single-instance execution to prevent data contention or race conditions. To allow overlapping executions, the 'max_concurrent_runs' parameter must be adjusted in the job settings.

This is a critical configuration for scenarios where independent data partitions need to be processed in parallel to meet aggressive latency requirements in a production system.

Exam trap

Candidates often look for code-level synchronization locks or cluster scaling options, missing that Databricks job settings explicitly block concurrent runs by default.

15
Multi-Selecthard

A data engineer is troubleshooting a Databricks job that fails with a 'TaskFailed' error. The job uses a cluster with autoscaling enabled. The engineer suspects that the failure is due to memory issues on the workers. Which TWO actions should the engineer take to diagnose and resolve the issue? (Choose two.)

Select 2 answers
A.Increase the driver node's memory to handle the failing tasks.
B.Increase the number of partitions when reading the data to reduce per-task memory usage.
C.Disable autoscaling to ensure the cluster has a fixed number of workers.
D.Check the cluster's event log for 'ExecutorLostFailure' events indicating out-of-memory errors.
E.Review the Spark UI for the job's run to identify stages with high garbage collection time or spill.
AnswersD, E

The cluster event log records executor losses and out-of-memory errors. 'ExecutorLostFailure' often indicates that an executor was killed due to memory limits. Reviewing the event log helps confirm if memory is the culprit and provides details such as which executor failed and why. This is a key diagnostic step.

Why this answer

To diagnose memory issues, the engineer should use the Spark UI to examine memory-related metrics such as garbage collection and spill, and check the cluster event log for executor out-of-memory errors. These steps confirm whether memory is the bottleneck. Once confirmed, the engineer can adjust cluster configuration, such as increasing executor memory or optimizing the job.

Disabling autoscaling or increasing driver memory does not address worker memory problems, and increasing partitions is a potential fix but not a primary diagnostic step.

Exam trap

The trap here is assuming that driver memory or autoscaling settings are the cause, when the issue is likely per-executor memory and requires inspection of Spark UI and event logs.

16
MCQhard

A data engineer is using Databricks Asset Bundles to deploy a job that runs a Python wheel task. The bundle is deployed to a production workspace using a service principal. The job fails with the error: `Library installation failed for library due to user error: Could not find wheel file`. The engineer confirms the wheel file exists in the bundle's `dist` folder. What is the most likely cause of this failure?

A.The wheel file path in the task definition is relative to the workspace root instead of the bundle root.
B.The wheel file is not compatible with the Databricks Runtime version used by the cluster.
C.The service principal lacks permission to read from the `dist` folder in the workspace.
D.The wheel file is not included in the bundle's artifact definition, so it is not uploaded to the workspace.
AnswerD

In Databricks Asset Bundles, artifacts such as Python wheels must be explicitly defined in the `artifacts` section of `databricks.yml`. If the wheel is not listed, it will not be uploaded during deployment, and the job cannot find it. The `dist` folder alone does not guarantee inclusion; the artifact must be declared and built.

Why this answer

For a Python wheel task in a Databricks Asset Bundle, the wheel must be declared as an artifact in the bundle configuration. This ensures it is built and uploaded to the workspace during deployment. Without this, the job cannot locate the wheel, leading to the error.

The other options incorrectly attribute the failure to permissions, path resolution, or compatibility.

Exam trap

The trap here is assuming that placing the wheel in the local `dist` folder is sufficient, when the bundle must explicitly declare and upload artifacts.

17
MCQhard

A data engineer is automating the deployment of Databricks assets using CI/CD. The pipeline fails because the 'databricks-cli' command cannot find the workspace. What is the most likely cause?

A.The Databricks CLI version on the build agent is incompatible with the server.
B.The runner does not have the DATABRICKS_HOST and DATABRICKS_TOKEN environment variables set.
C.The workspace API is currently disabled for security reasons.
D.The CI/CD runner is missing the required library dependencies like PySpark.
AnswerB

The Databricks CLI looks for these specific environment variables to authenticate with the workspace. If they are absent, the CLI cannot establish a connection, leading to an error indicating that it cannot reach or find the workspace. This is the most common cause of failures in automated pipelines.

Why this answer

Deployment failures in CI/CD often stem from authentication or environment variable misconfiguration. The Databricks CLI requires a properly configured profile or environment variables containing the host URL and token. When these are missing, the CLI fails to authenticate or locate the workspace, preventing the deployment of code.

Ensuring that the runner has access to these secure variables is foundational for a reliable DevOps lifecycle for Databricks infrastructure.

Exam trap

Candidates frequently assume CI/CD deployment failures are caused by syntax errors in configuration files rather than missing runtime environment variables.

18
MCQhard

A data engineer is debugging a Databricks job that reads from a Delta table and writes to another Delta table. The job occasionally fails with 'ConcurrentAppendException'. The engineer wants to minimize failures while maintaining data correctness. Which approach should the engineer take?

A.Disable Delta Lake transaction logging on the target table
B.Increase the job cluster's autoscaling maximum worker count
C.Use optimistic concurrency control with retry logic and partition the target table to reduce conflicts
D.Switch the target table to a Parquet table to avoid transaction conflicts
AnswerC

ConcurrentAppendException occurs when two writers attempt to add files to the same partition. Delta Lake uses optimistic concurrency; retrying the transaction and partitioning the target table to isolate writes reduces the chance of conflicting appends. This maintains correctness while improving success rates.

Why this answer

ConcurrentAppendException arises when multiple writers append to the same Delta table partition. Delta's optimistic concurrency control detects the conflict. Adding retry logic allows transient conflicts to resolve, and partitioning the target table by a key that separates writers reduces overlapping appends, preserving correctness while minimizing failures.

Exam trap

The trap here is thinking that scaling compute or changing file format solves a transactional concurrency conflict, when the fix lies in write isolation and retry behavior.

19
Multi-Selecthard

A data engineer is investigating a job failure that occurred only in the production environment. Which TWO features in Databricks help in comparing the production environment to the development environment?

Select 2 answers
A.Use the 'View as JSON' feature in the Jobs UI to compare job configurations.
B.Use the 'Cluster Logs' to compare the OS kernel versions of the nodes.
C.Review the git branch history and configuration files in the CI/CD pipeline.
D.Enable the 'Debug Mode' on the Spark Driver to see raw system calls.
E.Query the 'workspace_users' table to see who ran the job last.
AnswersA, C

Exporting and comparing job JSON configurations is an effective way to identify discrepancies in parameters, cluster sizes, or timeout settings. This allows engineers to systematically check for configuration drift between the development workspace and the production workspace, which is a common source of environment-specific failures.

Why this answer

Comparing environment configurations is essential for troubleshooting parity issues. Databricks provides tools like the Jobs UI for export and version control integration for code, which help engineers identify subtle differences between environments. Ensuring parity is critical for reducing 'it works on my machine' scenarios.

By utilizing these tools, engineers can verify cluster settings, library versions, and code versions to isolate why a process fails in production but succeeds in development.

Exam trap

Candidates often suggest manual code comparison or running both jobs simultaneously. They fail to see that configuration parity is best verified via the Jobs UI export and standardized CI/CD version control.

20
MCQmedium

A data engineer deploys a Databricks Job that runs a notebook task. The notebook writes to a Delta table in Unity Catalog. The job fails with the error: 'PERMISSION_DENIED: User does not have USE CATALOG on catalog 'prod'.' The engineer confirms the job's service principal has USE CATALOG granted on the catalog. Which configuration should the engineer check next?

A.The notebook's default language setting
B.Whether the job is configured to run as the service principal or as a different user
C.The job cluster's Spark config for spark.databricks.acl.enabled
D.The cluster's autoscaling min and max worker counts
AnswerB

Unity Catalog privileges are evaluated against the identity that executes the job. If the job's run-as setting points to a user or a different service principal than the one granted USE CATALOG, the error appears even though the intended principal has the privilege. Confirming and correcting the run-as identity ensures the granted privileges are applied to the execution context.

Why this answer

Unity Catalog privileges are enforced based on the identity that executes the job. When a service principal has USE CATALOG but the job runs as a different user or principal, the error still occurs. Checking the job's run-as configuration and aligning it with the granted identity resolves the mismatch and allows the notebook to access the catalog.

Exam trap

The trap here is assuming the error always means the grant is missing, when in fact the job may be running as a different identity than the one that was granted privileges.

21
Multi-Selectmedium

A data engineer is preparing to deploy a production Databricks workflow. Which TWO best practices should be implemented to ensure maintainability and robust error handling?

Select 2 answers
A.Hardcode credentials directly into the notebook for ease of access.
B.Use Databricks Git folders for version control of production code.
C.Configure email notifications for both job success and failure.
D.Avoid using libraries and rely only on built-in Spark functions.
E.Deploy code directly from the workspace to production.
AnswersB, C

Git folders provide essential version control functionality, allowing teams to manage code changes, handle merge requests, and maintain a history of deployments. This is fundamental for collaborative development and ensures that production code is peer-reviewed and consistent across different environments, preventing unauthorized or accidental changes.

Why this answer

Maintaining production environments requires strict adherence to modular code design and comprehensive observability. Using version control for notebooks and job configurations ensures that every change is tracked, audited, and reversible. Simultaneously, implementing robust notification alerts for job failures allows engineering teams to respond proactively to issues.

These two practices collectively minimize the risk of deployment errors and reduce the Mean Time to Resolution (MTTR) when unexpected failures occur in production pipelines.

Exam trap

Candidates frequently select 'manual code deployment' or 'cluster logging' instead of Git folders, mistakenly believing that basic dashboard monitoring is sufficient for production-grade code versioning and reliability.

22
Multi-Selectmedium

A data engineer is debugging a Databricks job that fails with a `SparkException: Job aborted due to stage failure` in production. They need to identify the root cause. Which two actions should they take to gather relevant diagnostic information? (Choose two.)

Select 2 answers
A.Review the Spark driver logs for the failed job run in the Databricks Jobs UI.
B.Run the job locally with a small sample of data to reproduce the error.
C.Examine the event log for the job cluster to see if nodes were terminated unexpectedly.
D.Check the cluster's init script logs to see if a library installation failed.
E.Inspect the Spark UI for the failed stage to identify skewed tasks or spills.
AnswersA, E

The Spark driver logs contain detailed error messages, including the stage that failed and the exception stack trace. This is the first place to look for root cause. The logs are accessible from the job run's detail page and provide context about the failure, such as data issues or resource problems.

Why this answer

To debug a Spark stage failure, the driver logs provide the exception details, and the Spark UI offers performance metrics that reveal issues like skew or spills. Together, they help identify whether the failure is due to data, code, or resource problems. The other actions are either for different error types or less direct for this specific failure.

Exam trap

The trap here is overlooking the Spark UI in favor of only logs, missing critical performance metrics that explain why the stage failed.

23
MCQmedium

Refer to the exhibit. A data engineer is deploying a production pipeline that references a table in the default schema. The job fails with the provided error. What is the root cause?

A.The cluster is running an outdated version of the Spark runtime.
B.The job is running under a service principal that lacks permissions to the Hive Metastore.
C.The table exists only in the temporary session catalog of a different notebook.
D.The cluster has insufficient memory to load the table metadata.
AnswerC

Temporary views or tables created in interactive notebook sessions are not persisted in the shared Hive Metastore and are scoped to the session. Since the job runs in a separate, isolated environment, it cannot access objects defined in the transient memory of a different interactive session.

Why this answer

The error indicates that the Spark session cannot resolve the table 'default.sales_data'. In Databricks, jobs often run in a different environment or context than an interactive notebook. If the table was created in an interactive session, it might not exist in the environment where the job runs, or the database context is missing.

This highlights the importance of using absolute paths or proper schema initialization in production code.

Exam trap

Test-takers often assume local temporary views or notebook session-scoped tables persist automatically when the code is deployed as a production job.

24
MCQhard

Refer to the exhibit. A Databricks job fails with a 403 Forbidden error when trying to write to the S3 bucket. Why does this happen?

A.The Databricks cluster needs to be restarted to apply the new IAM policy.
B.The policy is missing the 's3:PutObject' action required for writing data.
C.The S3 bucket policy is blocking the request despite the IAM policy.
D.The 'Resource' ARN is missing the suffix '/*'.
AnswerB

The provided policy explicitly only allows 's3:GetObject'. To successfully write data to an S3 bucket, the IAM policy must include 's3:PutObject' as well. The 403 Forbidden error happens because the service principal lacks the authorization to perform the write operation defined in the Spark job code.

Why this answer

The provided IAM policy only grants 's3:GetObject' permissions, which is read-only. For a job to write data, it needs additional permissions like 's3:PutObject'. In production, strict adherence to the principle of least privilege is required; however, the policy must also enable the necessary write operations for the task to complete successfully.

The 403 error is a direct consequence of this missing write-specific capability in the policy definition.

Exam trap

Candidates often look past explicit IAM action definitions, assuming read-only permissions like 's3:GetObject' are sufficient for writing files if storage bucket access is broadly enabled.

25
MCQmedium

A data engineer is debugging a slow-running query. They notice that the data is skewed, causing one task to take significantly longer than others. Which approach effectively addresses this skew?

A.Increase the cluster's disk size to allow for more local shuffle storage.
B.Add a salt column to the skewed key to distribute the data across more partitions.
C.Change the file format from Parquet to CSV to reduce overhead.
D.Enable 'Auto-scaling' to automatically add more nodes during the skewed stage.
AnswerB

Adding a salt to the join or grouping key forces Spark to distribute the data evenly across partitions. By breaking up the massive partition into smaller, manageable chunks, the workload becomes balanced, significantly reducing the execution time of the stage that was previously suffering from the data skew bottleneck.

Why this answer

Data skew occurs when one partition contains disproportionately more data than others, causing a single executor to bottleneck the entire stage. By using techniques like salt, broadcast joins, or repartitioning, the engineer can distribute the load more evenly across the cluster. This is essential for optimizing performance and preventing timeouts, ensuring that production jobs adhere to defined SLAs and utilize cluster resources efficiently.

Exam trap

Candidates often suggest simply increasing the cluster node count or core count, failing to realize that data skew leaves specific executors idle while one overloaded task bottlenecks the entire stage.

26
MCQmedium

A production Databricks workflow involves a task that runs a notebook. The notebook takes 15 minutes to finish, but the workflow is set to timeout after 10 minutes. What happens?

A.The workflow continues to run, but logs a warning about the duration.
B.The task is terminated and marked as 'Failed' by the scheduler.
C.The workflow pauses and waits for the engineer to manually approve continuation.
D.The task continues to run, but is moved to a background queue.
AnswerB

The scheduler enforces the timeout by killing the task process. Once terminated, the job workflow marks the task as 'Failed' because it did not complete successfully within the defined limits. This ensures that the system does not waste time and money on jobs that are behaving abnormally.

Why this answer

When a task exceeds its configured 'timeout' value, the Databricks scheduler forcibly terminates the task execution. This is a deliberate safety measure to prevent runaway processes from consuming cluster resources indefinitely. In production, this highlights the necessity of monitoring execution times and setting appropriate timeouts that account for normal data volume fluctuations while catching truly stuck jobs that could impact cost and resource availability.

Exam trap

Candidates often assume that a workflow timeout will gracefully cancel the notebook and mark it as 'Canceled' or 'Timed Out', overlooking that Databricks specifically categorizes exceeded task timeouts as 'Failed'.

27
MCQeasy

What is the primary benefit of using a Job Cluster instead of an All-Purpose Cluster for production workloads?

A.Job clusters support more advanced libraries than All-Purpose clusters.
B.Job clusters are more cost-effective and provide better workload isolation.
C.Job clusters can be manually resized while the job is currently running.
D.Job clusters are required to access data stored in the Unity Catalog.
AnswerB

Job clusters provide significant cost savings because they are transient and terminate automatically upon task completion. They also offer better workload isolation, as they are dedicated to a specific job, preventing resource contention or performance fluctuations caused by other users running interactive queries on the same cluster.

Why this answer

Job clusters are specifically designed for automated production workloads. They are cheaper because they are ephemeral, terminating as soon as the job finishes, and they provide better isolation, ensuring that production jobs do not interfere with other development tasks. This is a critical best practice for cost management and system reliability, ensuring that production pipelines run in a clean, predictable environment that scales according to actual need.

Exam trap

Candidates often choose All-Purpose clusters for production tasks because they stay alive, failing to recognize that Job clusters offer better workload isolation and significantly lower costs.

28
MCQhard

A data engineer is deploying a Databricks Asset Bundle (DAB) that defines a job with a notebook task. The bundle validates locally, but deployment fails with 'Error: cannot find notebook at path /Workspace/Users/dev@example.com/pipeline/ingest'. The engineer confirms the notebook exists in the workspace at that exact path. Which action should the engineer take to resolve the deployment failure?

A.Add the spark.databricks.workspace.path Spark config to the job cluster
B.Ensure the notebook path in the bundle configuration is relative to the bundle root and that the sync root includes the notebook
C.Convert the notebook task to a Python wheel task
D.Change the notebook task to use a Git source instead of a workspace path
AnswerB

DABs resolve notebook paths relative to the bundle root and synchronize files to the workspace. If the path is absolute or points outside the sync root, deployment cannot find it. Making the path relative and confirming the sync root includes the notebook ensures the file is uploaded and the job references the correct workspace location.

Why this answer

Databricks Asset Bundles synchronize files from the bundle root to the workspace. Notebook paths must be relative to the bundle root and included in the sync. An absolute or incorrectly rooted path causes deployment to fail even if the file exists in the workspace, because the bundle cannot map it to a synchronized file.

Exam trap

The trap here is assuming that because the notebook exists in the workspace, the bundle can reference it directly; bundles require relative paths within the sync root.

29
MCQeasy

A data engineer is configuring a Databricks Workflow that must run a notebook task only after a previous task that writes to a Delta table has completed successfully. The engineer wants to ensure that if the first task fails, the second task does not run. Which feature should the engineer use to define this dependency?

A.Configure the second task with a run_if condition set to ALL_SUCCESS.
B.Use a trigger to start the second task when the first task completes, using a file arrival trigger on the Delta table location.
C.Add a condition task between the two tasks that checks the Delta table for new records and only proceeds if records exist.
D.Set the depends_on property of the second task to reference the first task's task_key.
AnswerD

The depends_on property in a Workflow task definition specifies upstream tasks that must complete successfully before the dependent task starts. By referencing the first task's task_key, the second task will only run if the first succeeds, satisfying the requirement.

Why this answer

Task dependencies in Databricks Workflows are defined using the depends_on property, which lists upstream task keys. This ensures the dependent task runs only after the specified upstream tasks complete successfully. It is the fundamental way to build a directed acyclic graph (DAG) of tasks within a job.

Exam trap

The trap here is confusing run_if conditions with task dependencies; run_if controls behavior based on overall job status, not on specific upstream tasks.

30
MCQeasy

A data engineer needs to grant a group of users the ability to run a specific Databricks job but not modify its configuration. The job is managed by a service principal. Which permission level should be assigned to the group on the job?

A.Can View
B.Can Run
C.Can Manage
D.Can Attach To
AnswerB

Can Run allows users to trigger the job and view its runs but not edit its configuration. This exactly matches the requirement: the group can execute the job but cannot change its settings. It is the appropriate permission level for operators who need to run jobs without administrative privileges.

Why this answer

Databricks job permissions include Can View, Can Run, Can Manage, and sometimes Is Owner. Can Run permits users to execute the job and view run results but prevents them from editing the job's configuration. This aligns with the requirement to allow running without modification.

Can View is too restrictive, while Can Manage grants excessive privileges. Can Attach To is unrelated to job permissions.

Exam trap

The trap here is confusing job permissions with cluster permissions, or assuming that Can View also allows running.

31
Multi-Selectmedium

A data engineer is preparing to deploy a production Databricks Workflow that must be maintainable and auditable. The engineer wants to ensure that changes to the workflow are tracked and that failures can be diagnosed quickly. Which TWO practices should the engineer implement? (Choose two.)

Select 2 answers
A.Manually document each deployment in a shared spreadsheet
B.Use an all-purpose cluster instead of a job cluster for the workflow
C.Configure email notifications for job success and failure
D.Store the workflow definition in a Git repository and deploy using Databricks Asset Bundles
E.Enable job run logs to be delivered to a cloud storage location for long-term retention
AnswersD, E

Version-controlling workflow definitions in Git and deploying via Databricks Asset Bundles provides change tracking, code review, and reproducible deployments. This practice ensures that every modification is auditable and that the deployed workflow matches the reviewed source, which is essential for maintainability and compliance in production environments.

Why this answer

Version-controlling workflow definitions in Git and deploying with Databricks Asset Bundles ensures changes are tracked and reproducible. Delivering job logs to cloud storage retains diagnostic data for troubleshooting and auditing. Together, these practices provide the maintainability and auditability required for production workflows.

Exam trap

The trap here is choosing notification or manual documentation practices that seem helpful but do not provide the version control and durable logging needed for maintainability and auditing.

Ready to test yourself?

Try a timed practice session using only Debugging and Deploying questions.