Courseiva

Databricks-DE-Pro · domain

troubleshooting

Practise Databricks Certified Data Engineer Professional troubleshooting practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

267 questions42 easy137 medium88 hard

Focused practice

Practice troubleshooting questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about troubleshooting

troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common troubleshooting exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All troubleshooting questions (267)

Click any question to see the full explanation, or start a practice session above.

1

A data engineer is tasked with reducing compute costs for an interactive SQL analytics workspace that runs sporadic, highly unpredictable queries. The jobs experience cold start delays and occasional out-of-memory errors due to sudden concurrency spikes. Which TWO strategies should the engineer implement to balance cost efficiency and performance?

Hard
2

Refer to the exhibit. Which action is the most appropriate to resolve this memory-related failure during the job execution?

Hard
3

A data engineering team runs a nightly batch job on a Databricks job cluster. The job reads a large Parquet dataset, performs transformations, and writes the result to a Delta table. The cluster is configured with autoscaling from 4 to 16 workers and uses the default Spark configuration. The team observes that the job runs for 2 hours, but the cluster's CPU utilization is only around 30% throughout the run. They want to reduce cost without increasing runtime. Which action is most likely to achieve this?

Medium
4

Which of the following describes the correct behavior of Unity Catalog's 'Data Lineage' when used for security compliance?

Hard
5

A data engineer needs to ensure that a Databricks job can be retried automatically if it fails due to a transient cluster error. The job is configured with a maximum of 3 retries. After the third failure, the engineer wants to receive an email notification. Which feature should be used to accomplish this?

Easy
6

Refer to the exhibit. The engineer wants to replace only one specific partition in the 'orders' table. What is the best method in Databricks?

Medium
7

You are monitoring a long-running Databricks job. You notice that the memory usage on the driver node is steadily increasing until it crashes. Which debugging action is most appropriate?

Medium
8

What is the primary difference between sharing data via Delta Sharing compared to sharing data via Databricks-to-Databricks sharing?

Medium
9

You are building a pipeline and notice that the 'Gold' layer tables are experiencing significant write latency due to frequent small file commits. What is the most effective way to resolve this while maintaining ACID integrity?

Hard
10

A Data Engineer is using Delta Live Tables to process a stream of user events. The `user_id` column should be unique in the target table, but the source may contain duplicate events due to retries. The engineer wants to keep only the latest event for each `user_id` based on the `event_timestamp`. Which Delta Live Tables feature should be used?

Medium
11

A data engineer is using Lakehouse Federation to query data from an external PostgreSQL database. The engineer has created a foreign catalog and foreign schema. Which statement accurately describes how the data is accessed during a query?

Easy
12

Which of the following describes the purpose of the 'Gold' layer in a Lakehouse?

Medium
13

A Data Engineer needs to ensure that a notebook job is not consuming excessive costs. Which monitoring tool provides the best view of DBU consumption per job?

Medium
14

What is the primary advantage of using Delta Lake as the sink for your data ingestion pipelines compared to raw Parquet files?

Medium
15

A data engineer is responsible for a production Databricks SQL warehouse that serves multiple teams. The engineer needs to set up monitoring to detect when query performance degrades due to resource contention. Which two metrics should the engineer monitor to identify this issue? (Choose two.)

Medium
16

A Data Engineer needs to share a Delta table with a partner organization using Delta Sharing. The partner does not use Databricks. Which component must the Data Engineer generate to facilitate this secure connection?

Medium
17

A Data Engineer needs to monitor the health of a Delta Live Tables (DLT) pipeline. Which metric should they monitor to track the number of data quality violations over time?

Medium
18

You are ingesting data from multiple source systems with varying file formats (JSON, CSV, Parquet) into a centralized Bronze landing zone. Which architecture pattern is the most scalable for maintaining this ingestion layer?

Hard
19

When using Delta Sharing to share data with a recipient, what is the best way to handle updates to the shared data?

Medium
20

A data engineer has deployed a Databricks SQL dashboard that queries a gold-layer table. The dashboard is used by executives every morning. The engineer wants to be notified if the dashboard's underlying query fails or returns zero rows, which would indicate a data pipeline issue. Which Databricks feature should they use to set up this notification?

Easy
21

A company requires data to be physically deleted from the Bronze layer for GDPR compliance. What is the correct procedure to ensure complete removal?

Hard
22

A data engineer is building a streaming ingestion pipeline from Apache Kafka to a Delta table. The pipeline must perform deduplication on a unique event_id field and handle late-arriving data. The engineer wants to use Structured Streaming with a watermark of 10 minutes. Which of the following approaches correctly implements deduplication and watermarking?

Medium
23

Which statement correctly describes the relationship between Unity Catalog and Databricks SQL Warehouses when using Lakehouse Federation?

Medium
24

Which Databricks feature should be used to gain observability into access patterns and security events across the entire workspace?

Easy
25

A Data Engineer is using Delta Live Tables to process a streaming source that contains duplicate records based on an 'event_id'. The engineer needs to ensure that only the latest record for each 'event_id' is retained in the target table, and the pipeline should handle late-arriving data. Which DLT feature should be used?

Hard
26

Which capability is provided by Databricks' integration with cloud-native monitoring tools (e.g., CloudWatch, Azure Monitor)?

Medium
27

A Data Engineer is tasked with cleaning a dataset in Databricks. The dataset contains a column 'phone_number' with various formats, including parentheses, dashes, and spaces. The engineer needs to standardize all phone numbers to a digits-only format (e.g., '1234567890'). Which approach is most efficient and scalable?

Easy
28

A data engineer is managing a Delta Share that includes a table with frequent updates. The engineer wants to ensure that recipients always see the latest version of the data without manual intervention. Which statement is correct regarding how Delta Sharing handles updates to shared tables?

Hard
29

You are using the Databricks CLI to automate workspace tasks. Which THREE of the following statements correctly identify capabilities of the Databricks CLI?

Medium
30

A Data Engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data up to 2 hours and produce correct aggregations per 10-minute window. The engineer wants the streaming query to automatically clean up old state so the job does not accumulate unbounded state. Which combination of Structured Streaming features should be used?

Medium
31

A data engineer needs to audit which users have accessed a Unity Catalog table containing sensitive data. They want to see a record of all queries that read from the table over the past 30 days. Which Unity Catalog feature should they use?

Easy
32

A financial services firm maintains a Delta Lake table of account transactions that must support both current-state queries and full audit history of every change, including corrections that arrive days later. Regulators require the ability to query the table as it existed at any prior date. Which Delta Lake capability should the engineer rely on to satisfy the audit requirement?

Medium
33

A data engineer needs to restrict access to personally identifiable information (PII) columns in a Unity Catalog table for a group of analysts. Which Unity Catalog feature should be used to enforce this policy while ensuring data remains queryable?

Medium
34

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the ingestion must handle late-arriving data and produce correct aggregations. The engineer wants to ensure that watermarks are applied correctly. Which approach should be used?

Medium
35

A data engineer is building a streaming ingestion pipeline using Databricks Auto Loader to ingest JSON files from cloud storage into a Delta Bronze table. The pipeline must handle schema evolution without failing and must minimize the number of files that require reprocessing when the schema changes. The engineer wants to configure Auto Loader appropriately. Which two configuration settings should be used to achieve these requirements? (Choose two.)

Hard
36

You have a large Spark DataFrame that you need to filter and save as multiple smaller Parquet files based on the values in a 'region' column. Which method should you use to optimize the file layout for subsequent queries?

Medium
37

A Data Engineer needs to enforce a NOT NULL constraint on a specific column in a Delta table while maintaining the ability to perform high-performance streaming writes. Which approach is the most efficient and native method to ensure this data quality requirement?

Medium
38

A Data Engineer is implementing a medallion architecture. Which THREE steps are critical for effectively implementing a high-quality 'Silver' layer from 'Bronze' data?

Hard
39

Refer to the exhibit. Why is Task D marked as 'Skipped'?

Medium
40

When ingesting data using Auto Loader, what is the purpose of the 'cloudFiles.schemaLocation' parameter?

Medium
41

A data engineer manages a Delta table that is used for both batch analytics and frequent small updates from a streaming job. The table is not partitioned, and the engineer notices that queries are slowing down as the table grows. The engineer wants to improve query performance without changing the table schema or partitioning strategy. Which action should the engineer take?

Medium
42

A Data Engineer needs to perform an 'upsert' operation on a Delta table. Which command provides the most efficient way to merge new data with existing records in a single transactional step?

Medium
43

A Data Engineer is using Delta Live Tables to process a stream of financial transactions. The pipeline must ensure that each `transaction_id` appears only once in the target table, even if the source stream contains duplicates due to at-least-once ingestion. The engineer wants to use the `APPLY CHANGES` API. Which combination of settings will achieve this with minimal data loss?

Hard
44

Which TWO of the following are valid ways to pass configuration parameters to a Databricks Job task? (Select TWO)

Medium
45

A data engineer is configuring a Unity Catalog external location to allow a service principal to write to an ADLS Gen2 container. The storage credential uses a managed identity. The engineer grants the service principal `WRITE FILES` on the external location. However, when the service principal attempts to write, it fails with a permissions error. The engineer verifies that the managed identity has the Storage Blob Data Contributor role on the container. What is the most likely cause of the failure?

Hard
46

A data engineer is configuring audit logging for a Unity Catalog-enabled workspace. The security team wants to capture all access to tables and the granting of privileges. Which Databricks feature should the engineer enable to collect these audit events?

Easy
47

Refer to the exhibit. A data engineer creates an instance pool to reduce cluster startup times for development teams. However, finance reports indicate unexpected cloud infrastructure charges. Based on the configuration shown in the exhibit, what is the primary driver of these unexpected costs?

Hard
48

A data engineer is optimizing a PySpark job that processes a large DataFrame and writes the result to a Delta table. The job currently uses repartition(100) before writing, but the output consists of many small files. The engineer wants to reduce the number of output files without shuffling the entire dataset again. Which approach should be used?

Medium
49

A data team is using Liquid Clustering on a Delta table. How does this feature improve performance compared to traditional Z-Ordering or Partitioning?

Hard
50

You are managing a large-scale data lakehouse. You notice that your Spark jobs are frequently failing due to disk space issues on worker nodes. Which monitoring feature should you implement to proactively capture this trend?

Hard
51

A data engineer is designing an ETL pipeline processing high-frequency streaming data into Delta tables on Databricks. The pipeline experiences frequent small file creation and high metadata overhead, degrading query performance. Which optimization technique should the engineer implement to resolve this issue?

Medium
52

Refer to the exhibit. A Data Engineer is attempting to merge data into a table with these constraints defined. If the incoming batch contains rows that violate these rules, what is the default behavior of the Delta Lake engine during the merge operation?

Hard
53

A data engineer wants to ensure that all data in a specific catalog is encrypted at rest. Which feature should they verify is enabled within the Unity Catalog metastore configuration?

Medium
54

A data engineer is reviewing a Databricks job that runs on a job cluster and reads a large Delta table. The engineer notices that the job takes a long time to start because the cluster is provisioned from scratch each time. The engineer wants to reduce the startup time without increasing cost significantly. Which action should the engineer take?

Easy
55

A data engineer has a Unity Catalog managed table `sales.raw.transactions` that contains a column `customer_email` with PII. Analysts in the `marketing_analysts` group need to query the table for aggregate reporting but must never see individual email addresses. The engineer wants to enforce this dynamically without creating a separate view or copy of the data. Which Unity Catalog feature should the engineer use?

Medium
56

Your organization runs numerous batch data engineering pipelines using standard Databricks jobs. Finance reports indicate that compute costs are inflated due to cluster startup times and rigid over-provisioning. Which optimization approach provides the best balance of cost savings and execution reliability for scheduled production batch jobs?

Medium
57

Which TWO of the following are benefits of using the Databricks Delta Lake 'Optimize' command? (Select TWO)

Medium
58

A data engineer needs to ensure that PII data in a Delta table is accessible only to users in the 'HR_Manager' group. Which approach provides the most granular and scalable security implementation?

Medium
59

When designing a Data Quality framework in Databricks, what is the recommended approach for handling 'quarantined' records?

Medium
60

A Data Engineer needs to verify that the column 'user_id' is unique in a critical Gold table. What is the most efficient, non-blocking way to perform this check in a production environment?

Medium
61

When migrating an existing Hive metastore to Unity Catalog, what is the most important security consideration regarding object naming?

Medium
62

A data engineer is developing a PySpark job that reads from a Delta table and performs a series of transformations. The engineer notices that the job is slow and suspects that the query plan is not optimized because statistics are outdated. Which command should the engineer run to update the statistics for the Delta table to improve query performance?

Medium
63

A data engineer wants to share a Delta table with an external partner who does not have a Databricks account. The partner needs to access the data using Python. Which method should the engineer recommend to the partner for accessing the shared data?

Easy
64

Which Databricks feature provides the most granular view of data quality metrics over time for a Delta Live Tables pipeline?

Medium
65

A Data Engineer is using Databricks Asset Bundles (DABs) to manage a project. Where should the engineer define the job settings, such as clusters, schedules, and task dependencies?

Medium
66

A data engineer needs to inspect the logs of a long-running Databricks job that has already completed. Where should they navigate in the Databricks UI to find the driver logs?

Easy
67

Which THREE actions can help reduce the 'shuffle' operations in a Spark job?

Hard
68

A data engineer is using PySpark to cleanse a DataFrame containing customer addresses. The 'zip_code' column has some values with leading zeros that were stripped during CSV ingestion. The engineer needs to restore all zip codes to a fixed 5-character length by padding with leading zeros. Which function should be used?

Easy
69

Refer to the exhibit. A user encounters this error when running a query. What is the correct action to resolve this issue while maintaining the security model?

Hard
70

What is the primary benefit of the Medallion architecture in a Databricks Lakehouse?

Easy
71

A data engineer is deploying a Databricks job using Databricks Asset Bundles. They want to ensure that the job uses a specific cluster configuration that is defined once and reused across multiple tasks. Which bundle feature should they use?

Easy
72

Which TWO of the following are primary benefits of using Unity Catalog for data governance in Databricks?

Medium
73

A data engineer is configuring monitoring for a production Databricks cluster. Which TWO metrics are best suited to identify potential performance bottlenecks related to worker nodes?

Medium
74

A data engineer is using Databricks Repos to manage a project. They need to ensure that the production job always uses the code from the 'main' branch, even if developers push changes to other branches. Which Git reference should be used in the job configuration to achieve this?

Medium
75

A data engineer is ingesting data from an Apache Kafka topic into a Delta Lake table using Structured Streaming. The Kafka topic receives messages with a timestamp field in the value payload, but the messages can arrive out of order by up to 10 minutes. The engineer wants to perform time-windowed aggregations on the ingested data while minimizing state store overhead. Which approach should be used to handle the out-of-order data correctly?

Hard
76

Refer to the exhibit. An administrator reviews the cluster configuration JSON for an all-purpose interactive development cluster used by data engineers. Based on Databricks cost and performance optimization best practices, which specific parameter in this configuration represents the highest risk for unnecessary financial expenditure?

Medium
77

An engineer is building a Gold-layer star schema for a sales analytics workload. The business wants to analyze revenue by product, by store, and by promotion independently, and also drill down through a hierarchy of region to country to city. Which dimensional modeling structure best supports these requirements?

Easy
78

A data engineer is ingesting streaming data from Apache Kafka into a Delta table using Databricks Structured Streaming. The engineer wants to ensure exactly-once processing and handle late-arriving data. Which combination of features should the engineer use?

Medium
79

A data engineer maintains a Delta Lake table that stores 5 years of order data. Analysts frequently query the most recent 90 days, but compliance requires that older data remain queryable. The table is currently partitioned by order_date and has 200,000 small files because data arrives continuously via Structured Streaming. Queries on the last 90 days are slow and expensive. Which combination of actions will most effectively reduce query cost and improve performance for the recent-data queries?

Hard
80

A Databricks workspace has a Delta table 'transactions' partitioned by 'txn_date'. Analysts frequently run queries that filter on 'txn_date' but also occasionally filter on 'account_id' alone. The table has 10 TB of data, and the team wants to improve performance for the 'account_id' queries without changing the partitioning scheme. Which Delta feature should they implement?

Medium
81

A data engineer is using Delta Lake to manage a table that receives frequent updates and deletes. The engineer notices that query performance has degraded over time due to many small files. Which command should be used to optimize the table by compacting small files and improving query performance?

Hard
82

A retail company uses a Databricks Lakehouse with a star schema in the Gold layer. Their fact_sales table has billions of rows and is partitioned by sale_date. Analysts frequently run queries that filter on product_id and join to dim_product. Currently, queries scanning the entire fact table are slow. To improve performance for these queries, which approach is most appropriate?

Medium
83

A data engineer is using Databricks Repos to manage code for a production job. They need to ensure that the job always uses the latest version of the code from a specific branch. Which Git operation should they perform before running the job?

Easy
84

A financial institution uses a Databricks Lakehouse with a Silver table transactions that is partitioned by transaction_date. The table is frequently queried with filters on transaction_date and account_id. The data engineering team notices that queries filtering on account_id are slow because they scan all partitions. They want to optimize the table to accelerate these queries without repartitioning. Which Delta Lake feature should they use?

Hard
85

Which THREE techniques are recommended for improving the performance of Spark SQL joins on large Databricks tables?

Medium
86

You are writing a Spark application that uses a Broadcast Hash Join. You want to force the join to use broadcast optimization for a specific table. How do you implement this in PySpark?

Medium
87

Which of the following is the best practice for managing alerts for a mission-critical production pipeline?

Easy
88

A retail company is designing a Gold-layer dimension table in Delta Lake for its product catalog. The catalog changes slowly: a product's category is occasionally reclassified, but historical sales fact rows must continue to reflect the category that was valid at the time of each sale. The team wants to avoid duplicating the entire product row for every change. Which Delta Lake modeling technique should the engineer implement?

Medium
89

A data engineer is troubleshooting a Databricks Workflow where a downstream task relies on an upstream task's output. Which TWO actions ensure the data dependency is correctly handled during a failure scenario?

Hard
90

An engineer needs to optimize a massive table that is frequently joined with other large tables. Which strategy is most effective for performance?

Hard
91

A data scientist reports that their notebook takes 20 minutes to initialize, even when the cluster is running. What is the most likely reason for this high initialization time?

Medium
92

A data engineer is reviewing the event log of a Databricks job that has just failed. They need to determine the exact cause of the failure. Which event type in the event log indicates that a task failed due to an exception?

Medium
93

You are optimizing a PySpark job that reads from a Delta table. You notice skewed data distribution on the 'customer_id' column, causing Task-level stragglers. Which transformation should you apply to the DataFrame to mitigate this skew during a join operation?

Medium
94

Which design pattern is best suited for handling late-arriving data in a medallion architecture?

Medium
95

A data engineer is troubleshooting a Databricks job that fails with a `SparkException: Job aborted due to stage failure` and the error log shows `java.lang.OutOfMemoryError: GC overhead limit exceeded` on an executor. The job processes a large dataset using a `groupByKey` operation. Which action should the engineer take to resolve the issue while minimizing changes to the existing code?

Hard
96

A data engineer needs to grant a service principal permission to read data from a Unity Catalog table named sales.orders. The service principal is used by an automated job and should have only the minimum necessary privileges. Which Unity Catalog privilege should be granted on the table to allow the service principal to read data?

Easy
97

A data engineer is configuring a Databricks workspace with Unity Catalog. The security team wants to ensure that all data access is logged for compliance auditing. The engineer enables audit logs and configures delivery to a cloud storage location. Which Unity Catalog object should the engineer use to query audit logs for data access events?

Easy
98

A data engineer notices that a Databricks SQL warehouse used for executive dashboards runs 24/7 but is only actively queried during business hours. The warehouse is a Pro-sized warehouse with auto-stop set to 10 minutes. The team wants to reduce cost without affecting dashboard availability during business hours. Which action should the data engineer take?

Easy
99

A logistics company wants to analyze shipment delays. The fact table `fact_shipments` has a `delay_minutes` measure. The team needs to slice delays by the reason for delay, which can be one of several predefined categories. Which dimension modeling approach is most suitable?

Medium
100

A team is building a streaming pipeline that processes millions of events per second. They are using Structured Streaming with a Delta Lake sink. What is the most effective way to optimize the performance and cost of this write-heavy workload?

Medium
101

Which of the following describes the behavior of the 'OPTIMIZE' command in Databricks?

Medium
102

A data engineer is deploying a Databricks job that uses a Python wheel task. The job fails with the error: 'ModuleNotFoundError: No module named 'my_library''. The wheel file is stored in DBFS at 'dbfs:/FileStore/wheels/my_library-0.1.0-py3-none-any.whl'. The job cluster is configured with a cluster policy that restricts library installation from DBFS. What is the most likely cause of the failure?

Hard
103

In the medallion architecture, which layer is primarily responsible for applying business logic and historical aggregations?

Easy
104

A data engineer is designing a Gold layer table for a retail company. The table must support efficient queries that filter on product_category (low cardinality) and sort by transaction_timestamp (high cardinality). The table is expected to grow to petabytes. Which Delta Lake table design should the engineer choose to optimize both filtering and sorting?

Medium
105

You are developing a Delta Live Tables (DLT) pipeline and need to ensure high data quality. Which TWO of the following statements correctly describe how Expectations work within DLT?

Hard
106

Which THREE strategies should a data engineer use to optimize the debugging of failed production Databricks Jobs?

Hard
107

An organization wants to restrict data access to only allow connections from specific corporate IP ranges. Which Databricks feature should be configured to implement this network security requirement?

Medium
108

A financial services firm stores market data in an external location registered in Unity Catalog as `s3://firm-market-data/`. The security team requires that only a specific IAM role, assumed by a Unity Catalog storage credential, can read the bucket, and that no Databricks user can bypass Unity Catalog to read the data directly with their own cloud credentials. The Data Engineer must configure the storage credential. Which configuration achieves this?

Hard
109

Which TWO of the following are benefits of using Unity Catalog for managing data governance in a multi-workspace environment?

Medium
110

You need to perform a deduplication task on a streaming source that includes late-arriving data. Which Delta Lake feature is best suited to manage this while ensuring efficient state cleanup?

Hard
111

A data engineer is configuring a Unity Catalog storage credential to access an AWS S3 bucket. The organization's security policy requires that Databricks assumes an IAM role, and that no long-lived AWS access keys are stored in Databricks. The engineer has created an IAM role with a trust policy and an external ID. Which action must the engineer take to complete the storage credential configuration in Unity Catalog?

Hard
112

A Data Engineer is setting up Lakehouse Federation for a PostgreSQL database. Which TWO steps are required to ensure that users can securely query the data using Unity Catalog?

Medium
113

Which THREE actions are required to properly implement a secure data sharing strategy using Delta Sharing?

Hard
114

A data engineer is setting up a Delta Share to provide a partner with access to a specific table. The partner uses Databricks and wants to query the shared data using their own Databricks workspace. What is the correct sequence of actions for the data engineer to enable this Databricks-to-Databricks sharing?

Easy
115

A data engineer is tuning a Spark job that reads from a Delta table and writes to another Delta table. The job uses a groupByKey operation followed by an aggregation. The engineer notices that the job is spilling to disk during the shuffle and taking a long time. The engineer wants to reduce shuffle spill and improve performance. Which action is most likely to help?

Hard
116

Which THREE of the following are benefits of using Delta Lake over standard Parquet files in Databricks?

Medium
117

A data engineer is using Databricks Auto Loader to ingest CSV files into a Delta table. The engineer notices that some files have a different delimiter (semicolon instead of comma). Which option should be used to handle this variation?

Easy
118

A data engineer is designing a Gold layer table that must support slowly changing dimension (SCD) Type 2 for a customer dimension. The source data arrives daily with updates to customer attributes. The engineer wants to implement this using Delta Lake. Which two features or techniques are essential for maintaining SCD Type 2? (Choose two.)

Hard
119

A retail company wants to analyze sales by product, store, and date. The data team is designing the Gold layer and needs to choose between a star schema and a snowflake schema. Which factor most strongly favors a star schema in a Databricks Lakehouse?

Easy
120

A data engineer at a healthcare company needs to share a Delta table containing patient records with an external research partner. The partner must only see aggregated statistics, not individual patient rows. The engineer wants to enforce this at the data sharing layer without creating a separate physical copy of the table. Which approach should the engineer take?

Medium
121

A data engineer wants to monitor the health of Delta Live Tables (DLT) pipelines and be alerted if a pipeline fails. Which approach is the most efficient and native way to achieve this?

Medium
122

A healthcare company uses a Databricks Lakehouse. The Silver layer contains a table patient_visits that is updated with late-arriving data. The table is partitioned by visit_date. The data engineering team needs to efficiently merge new data that may include updates to existing records and inserts of new records. They want to minimize the impact on existing data and ensure ACID compliance. Which Delta Lake operation should they use?

Hard
123

A data engineer needs to receive an email notification when a Databricks job fails. The job is scheduled to run every hour. The engineer wants to configure this notification with minimal effort and without writing additional code. Which approach should the engineer use?

Easy
124

A data engineer needs to grant a new data analyst the ability to query tables in the `sales` catalog, which is in Unity Catalog. The analyst should only be able to read data and not modify any tables or metadata. Which sequence of privileges should the engineer grant to the analyst?

Medium
125

A data engineer is tuning a Spark job and decides to use 'Z-Ordering' on a Delta table. Which THREE of the following are valid considerations when selecting columns for Z-Ordering?

Hard
126

A data engineer is implementing a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The engineer needs to ensure that the job can recover from failures without data loss or duplication. The job uses foreachBatch to perform upserts into the Delta table. Which checkpointing configuration is required to achieve exactly-once semantics?

Hard
127

A data engineer is building a Silver layer table that combines data from multiple Bronze tables. The engineer wants to ensure that the Silver table only contains the most recent version of each record based on a 'last_updated' timestamp. Which Delta Lake operation should be used to achieve this?

Easy
128

Refer to the exhibit. A data engineer executes this command in a Unity Catalog-enabled workspace. What is the immediate effect on the 'analyst_group'?

Hard
129

Where can a data engineer find the standard output and error logs for a specific task within a Databricks Workflow?

Easy
130

A data engineer is designing a pipeline and notices that the cost of processing is unexpectedly high during development. Which action provides the most immediate cost reduction when using Databricks?

Easy
131

When refining data in a Medallion architecture, why is it recommended to perform schema enforcement as early as possible in the Bronze layer?

Medium
132

A data engineer is building a Structured Streaming job that reads from a Kafka topic and writes to a Delta table. The pipeline must tolerate late-arriving data up to 10 minutes and update aggregations accordingly. The engineer wants to use a watermark on the event-time column. Which code snippet correctly applies the watermark and performs a 5-minute tumbling window aggregation?

Medium
133

A data engineering team maintains a Unity Catalog metastore in a Databricks workspace. They need to provide an external partner with read-only access to a specific Delta table, but the partner's analytics platform is not Databricks and does not support the Delta Lake protocol. The partner can consume Parquet files over a REST API. Which Unity Catalog feature should the team use to share the table?

Medium
134

Refer to the exhibit. Which configuration change is required to enable multiple instances of this job to run simultaneously?

Hard
135

A data engineer is troubleshooting a Databricks job that fails with a 'TaskFailed' error. The job uses a cluster with autoscaling enabled. The engineer suspects that the failure is due to memory issues on the workers. Which TWO actions should the engineer take to diagnose and resolve the issue? (Choose two.)

Hard
136

Which property should be configured to allow Databricks to automatically optimize the size of files during write operations in Delta Lake?

Hard
137

When partitioning a large Delta table, what is the best practice regarding the number of unique values in the partition column?

Easy
138

A data engineer is responsible for monitoring a production Databricks job that runs critical ETL tasks. The job occasionally fails due to transient issues such as cloud storage throttling or network timeouts. The engineer wants to set up automated alerts that notify the team only when the job fails after all retries are exhausted. Which TWO actions should the engineer take to achieve this? (Choose two.)

Hard
139

A data engineer is troubleshooting a Delta Live Tables pipeline that intermittently fails with 'StreamingQueryException: Job aborted due to stage failure'. The pipeline processes streaming data from a Kafka source. Which monitoring approach will best help identify the root cause of these intermittent failures?

Hard
140

A Data Engineer is using Delta Live Tables (DLT) to build a pipeline that ingests JSON files from cloud storage. The engineer defines a streaming table with expectations to enforce data quality. The expectation `@dlt.expect_or_drop("valid_timestamp", "timestamp IS NOT NULL")` is applied. During a pipeline run, 5% of records have a NULL timestamp. What is the outcome for those records, and how does it affect the pipeline?

Hard
141

A data engineer is using Auto Loader to ingest JSON files from cloud storage into a Delta table. The files contain a nested field 'address' with subfields 'city' and 'zip'. The engineer wants to flatten the nested structure during ingestion so that 'city' and 'zip' become top-level columns in the Bronze table. Which Auto Loader feature should be used to achieve this?

Easy
142

Which TWO of the following are primary benefits of using Unity Catalog for managing data lineage in Databricks?

Medium
143

When migrating to Unity Catalog, what is the best practice for managing existing data access permissions?

Medium
144

A data engineer manages a Databricks SQL warehouse that serves a dashboard used by the finance team. The dashboard queries have become slow during peak hours, and the engineer suspects that some queries are scanning excessive data. Which system table should the engineer query to analyze query performance and identify expensive queries?

Medium
145

A data engineer is responsible for a Delta Live Tables pipeline that ingests streaming data from multiple sources. The pipeline occasionally experiences delays, and the engineer needs to monitor the pipeline's health. Which two metrics should the engineer monitor to detect ingestion backlog and processing latency? (Choose two.)

Hard
146

Which approach is most appropriate for ingesting data from a JDBC source into Delta Lake where the source table has no 'updated_at' or 'version' column for incremental loading?

Medium
147

A Databricks SQL warehouse is experiencing high costs due to idle resources. Which TWO configurations should be implemented to effectively manage and reduce warehouse costs?

Hard
148

Which of the following describes the correct use of a UDF (User Defined Function) in PySpark for production pipelines?

Medium
149

A data engineer at a healthcare company needs to share a Delta table containing patient records with an external research partner. The partner uses a non-Databricks platform and requires read-only access to the latest data, with updates reflected in near real-time. The engineer creates a Delta Share and adds the table. Which additional step is required to allow the partner to access the shared data?

Medium
150

A data engineer is using Lakehouse Federation to query an external PostgreSQL database. The engineer creates a connection with the PostgreSQL JDBC URL and credentials, and then creates a foreign catalog. Users report that queries against foreign tables are slow and sometimes fail with connection timeouts. The engineer checks the connection and confirms the credentials are correct. What is the most likely cause of the performance and timeout issues?

Medium
151

A data engineer wants to monitor the performance of a Databricks cluster by tracking the average CPU utilization over time. Which Databricks feature should they use to visualize this metric?

Easy
152

A data engineering team experiences massive compute waste because interactive development notebooks are frequently left running overnight by engineers. As a Databricks administrator, which configuration should you implement at the cluster policy level to automatically mitigate this financial exposure without disrupting ongoing development work?

Medium
153

A data engineer has set up a Databricks SQL alert on a query that returns the count of failed jobs in the last hour. The alert is configured to trigger when the count exceeds 5. The engineer wants to receive notifications via email and also wants to view the alert history to understand past triggers. Which statement accurately describes the alert notification and history capabilities?

Medium
154

A data engineer is using Databricks Asset Bundles to deploy a job that runs a Python wheel task. The bundle is deployed to a production workspace using a service principal. The job fails with the error: `Library installation failed for library due to user error: Could not find wheel file`. The engineer confirms the wheel file exists in the bundle's `dist` folder. What is the most likely cause of this failure?

Hard
155

You are tasked with handling PII (Personally Identifiable Information) in your data pipeline. Which approach is best for protecting this data while maintaining the ability to perform analytics?

Medium
156

You are debugging a PySpark job that is experiencing severe memory pressure during a join on a massive column. You suspect data skew. Which approach is best to mitigate this issue?

Hard
157

A data engineer is automating the deployment of Databricks assets using CI/CD. The pipeline fails because the 'databricks-cli' command cannot find the workspace. What is the most likely cause?

Hard
158

Which THREE features are provided by Delta Lake when compared to standard Parquet files? (Select THREE)

Medium
159

A data engineer is using PySpark to process a large DataFrame and needs to reduce the number of partitions before writing to a Delta table to avoid creating too many small files. The DataFrame currently has 2000 partitions, each about 10 MB. The engineer wants to reduce the number of partitions to approximately 200 while minimizing data shuffling. Which approach is most appropriate?

Hard
160

A data engineer is using Delta Live Tables (DLT) to create a pipeline that processes streaming data. They need to ensure that the pipeline only processes new data since the last run and that the pipeline can recover from failures without reprocessing all data. Which combination of features should they use to achieve this?

Hard
161

A data engineer is building a Structured Streaming pipeline that reads from a Kafka topic and writes to a Delta table. The pipeline must handle late-arriving data and update already processed aggregates. Which watermark strategy should be used to allow updates to aggregates while bounding state store growth?

Medium
162

Which TWO statements regarding the use of 'APPLY CHANGES INTO' in Delta Live Tables (DLT) are correct?

Hard
163

Refer to the exhibit. An engineer observes that queries filtering on 'customer_id' are running slowly despite Z-Ordering. What is the most likely cause?

Medium
164

A data engineer is debugging a Databricks job that reads from a Delta table and writes to another Delta table. The job occasionally fails with 'ConcurrentAppendException'. The engineer wants to minimize failures while maintaining data correctness. Which approach should the engineer take?

Hard
165

A data engineer is using Delta Live Tables (DLT) to build a pipeline that ingests data from a streaming source. The pipeline must ensure that the target table is updated incrementally and that data quality constraints are enforced. The engineer wants to use expectations to drop invalid records while maintaining pipeline performance. Which TWO of the following are true regarding DLT expectations and their behavior? (Choose two.)

Hard
166

An engineer notices that a SQL warehouse is frequently hitting 'Max Concurrency' limits. Which log should they consult to identify which specific queries are consuming most of the warehouse resources?

Hard
167

You are migrating a legacy ETL process to Delta Live Tables (DLT). You have an existing table defined with a complex transformation that involves a custom Python function using a third-party library. How should you structure this in DLT to ensure the function is available and correctly applied?

Hard
168

When using the Unity Catalog, how should an engineer properly reference a table named 'sales' located in the 'finance' schema within the 'prod_catalog' catalog using Spark SQL?

Hard
169

What is the primary function of a 'Personal Access Token' (PAT) in Databricks, and why is it considered a security risk if not managed properly?

Medium
170

Refer to the exhibit. A Databricks administrator wants to restrict access to a specific SQL Alert. Based on the JSON policy, which statement accurately describes the current permission model for this alert?

Hard
171

Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?

Medium
172

A data engineering team is running a nightly batch job on a Databricks job cluster that processes a 10 TB Delta table. The job reads the entire table, performs transformations, and writes results to another Delta table. The team notices that the job takes 4 hours and consumes significant DBUs. They want to reduce runtime and cost without changing the business logic. The table is partitioned by ingestion date, but queries often filter on a high-cardinality column 'customer_id'. Which optimization technique is most appropriate to improve performance and reduce cost?

Medium
173

You are processing sensitive PII data in a Delta table. You need to ensure that specific columns containing PII are not readable by general data analysts while maintaining the ability to perform aggregate analysis on those rows. Which feature should you implement?

Medium
174

A data engineer is configuring a Unity Catalog external location to allow access to an S3 bucket. The security team requires that all access to the bucket be authenticated using a specific IAM role, and that the credentials not be stored in Databricks. Which Unity Catalog object should the engineer create to meet this requirement?

Medium
175

Which command should be used to display the history of transactions performed on a Delta table, including operations like overwrites and updates?

Easy
176

A data engineer is implementing column-level masking in Unity Catalog. They need to mask the 'email' column in the table 'prod.customers' such that only members of the 'hr_group' see the full email, while all other users see a masked version. The engineer creates a masking function and applies it using ALTER TABLE. Which statement correctly applies the mask?

Hard
177

Which metric should a data engineer prioritize when investigating a slow-running query in the Databricks SQL query history?

Easy
178

A data engineer is investigating a job failure that occurred only in the production environment. Which TWO features in Databricks help in comparing the production environment to the development environment?

Hard
179

Which property must be set to ensure a Spark Structured Streaming query can handle changes to the source data schema, such as adding a new column?

Hard
180

A data engineer is optimizing a Delta Lake table that experiences high read latency due to many small files. Which command should be executed to physically reorganize the data layout to improve query performance?

Medium
181

A data engineer deploys a Databricks Job that runs a notebook task. The notebook writes to a Delta table in Unity Catalog. The job fails with the error: 'PERMISSION_DENIED: User does not have USE CATALOG on catalog 'prod'.' The engineer confirms the job's service principal has USE CATALOG granted on the catalog. Which configuration should the engineer check next?

Medium
182

A data engineer is configuring a Delta Live Tables (DLT) pipeline that processes streaming data from Apache Kafka. The pipeline performs a series of transformations and writes to a Delta table. The engineer notices that the pipeline is experiencing high latency and wants to optimize it for cost and performance. The pipeline is set to continuous mode. Which configuration change is most effective to reduce cost while maintaining acceptable latency?

Hard
183

Refer to the exhibit. You are using Auto Loader to ingest data with evolving schemas. After running the job for a week, you realize that new columns added to the source JSON are not being captured in the destination table. What must you add to the configuration?

Hard
184

Which THREE of the following are essential components of an effective ingestion monitoring strategy in Databricks?

Medium
185

A data engineer is preparing to deploy a production Databricks workflow. Which TWO best practices should be implemented to ensure maintainability and robust error handling?

Medium
186

A data engineer supports a Delta Live Tables pipeline that ingests streaming data from Kafka. The pipeline sometimes experiences latency spikes, and the engineer needs to determine whether the bottleneck is in the ingestion stage or in downstream transformations. They want to use built-in observability without adding external tooling. Which approach provides the most direct insight into per-stage event processing times within the DLT pipeline?

Medium
187

When troubleshooting a job that frequently crashes due to 'Out of Memory' (OOM) errors, which TWO metrics or logs should be analyzed?

Hard
188

A Data Engineer is implementing column-level security on a Unity Catalog table `sales.customers` that contains `email`, `ssn`, and `region` columns. The requirement is that analysts in the `analyst` group see only the last four digits of `ssn` and a hashed `email`, while members of the `compliance` group see full values. The engineer plans to use column masks. Which TWO actions are required to meet the requirement? (Choose two.)

Medium
189

A data engineer needs to read a CSV file from cloud storage into a Spark DataFrame in Databricks. The file has a header row and uses commas as delimiters. The engineer wants to infer the schema automatically. Which code snippet correctly reads the file?

Easy
190

An organization requires that all data stored in their S3 bucket used by Databricks be encrypted using a Customer Managed Key (CMK). Which configuration must be performed to meet this requirement?

Medium
191

A data engineer is asked to implement column-level masking for a Unity Catalog table `main.hr.employees` that contains a column `ssn` with Social Security numbers. The requirement is that only members of the `hr_group` should see the full SSN, while all other users should see only the last four digits (e.g., XXX-XX-1234). The engineer decides to use a column mask function. Which statement accurately describes how to apply the mask?

Easy
192

What is the primary benefit of using Unity Catalog for data federation?

Medium
193

Your team is using a shared cluster for development. A user reports that their job is slow because the cluster memory is frequently filled by large data broadcasts. What configuration adjustment should you make to prevent this issue across all jobs on the cluster?

Hard
194

When running a PySpark job, you receive an 'Out of Memory (OOM)' error during a shuffle operation. Which configuration is the most appropriate to address this first?

Hard
195

A data engineer is ingesting data from an Azure SQL Database into a Delta Lake table using the JDBC connector in a Databricks notebook. The source table contains millions of rows, and the engineer wants to optimize the ingestion by reading the data in parallel. The source table has a numeric primary key column named 'id' that is evenly distributed. Which approach should the engineer use to achieve parallel reads?

Medium
196

A data engineer is using Lakehouse Federation to query a PostgreSQL database. The engineer notices that a query filtering on a column with a high cardinality is performing poorly, even though the remote database has an index on that column. What is the most likely reason for the poor performance?

Hard
197

A data engineer is debugging a Databricks job that fails with a `SparkException: Job aborted due to stage failure` in production. They need to identify the root cause. Which two actions should they take to gather relevant diagnostic information? (Choose two.)

Medium
198

A data engineer is using PySpark to cleanse a large dataset of customer records. The DataFrame `df` contains a string column `phone` with values like '123-456-7890', '(123) 456-7890', and '1234567890'. The engineer needs to standardize these to digits only (e.g., '1234567890'). Which transformation should be used?

Medium
199

When designing a production-grade data pipeline in Databricks, what is the recommended approach for managing secrets such as database credentials?

Medium
200

A data engineering team is modeling a large Delta Lake fact table that stores clickstream events. Analysts frequently run queries that filter by event_date and then aggregate by user_id, and the table receives continuous appends plus occasional late-arriving corrections. The team wants to reduce bytes scanned and improve join performance. Which two design choices are most appropriate? (Choose two.)

Hard
201

A data engineer is reviewing a Databricks job that runs a notebook to process a large Delta table. The job takes 45 minutes, and the engineer notices that the cluster spends a significant amount of time in the 'Pending' state before execution begins. The cluster is a job cluster with autoscaling enabled and no cluster pool. The engineer wants to reduce the overall job duration and cost. Which action should the data engineer take?

Medium
202

Refer to the exhibit. A data engineer is deploying a production pipeline that references a table in the default schema. The job fails with the provided error. What is the root cause?

Medium
203

A data engineer is developing a PySpark job that reads a large Delta table, performs a groupBy on a high-cardinality column, and writes the result to another Delta table. The job is experiencing performance issues due to data skew. The engineer wants to optimize the shuffle by using salting. Which approach correctly implements salting to distribute the skewed keys evenly?

Medium
204

A data engineer wants to monitor the health of a Delta Live Tables pipeline and receive alerts when the pipeline fails to meet its data quality expectations. The pipeline has several expectations defined. Which Databricks feature should the engineer use to set up these alerts?

Easy
205

A data engineer is implementing fine-grained access control on a Delta table in Unity Catalog that contains sensitive customer data. The requirement is to mask the `credit_card` column for all users except members of the `finance` group, and to filter out rows where the `region` column is not in the user's allowed regions. Which two Unity Catalog features should the engineer use? (Choose two.)

Hard
206

A data engineer is configuring a Databricks SQL warehouse to handle a workload that consists of many concurrent short queries during business hours and almost no queries at night. The engineer wants to minimize cost while ensuring low latency during peak hours. Which configuration should the engineer use?

Medium
207

A Data Engineer needs to encrypt data at rest within a Databricks workspace that uses a customer-managed key (CMK). What is the primary purpose of this configuration?

Hard
208

Refer to the exhibit. An engineer has configured the cluster settings as shown. What is the expected impact on the Delta table's performance and write operations?

Medium
209

A Data Engineer wants to monitor data quality trends over time for a critical table. Which tool provides the most native and easy-to-use visualization of these metrics?

Medium
210

An organization needs to share a dataset with a client who does not use Databricks. What is the most efficient and secure way to share this data using Unity Catalog?

Medium
211

You are performing a complex data transformation involving a self-join on a large, skewed table. Which technique is most effective for preventing data skew and improving join performance?

Hard
212

You are designing an ingestion pipeline that must handle massive bursts of data at irregular intervals. Which feature should you prioritize to ensure the ingestion process remains cost-effective?

Medium
213

A data engineer is reviewing the cost of a Databricks job that runs on a daily basis. The job uses an all-purpose cluster that is manually started and stopped by the engineer. The job typically runs for 30 minutes, but the engineer often forgets to stop the cluster, leading to hours of idle time. Which action should the engineer take to reduce cost?

Easy
214

When ingesting data from a Kafka topic into Delta Lake, what is the best way to handle out-of-order data arriving in the stream?

Medium
215

A data engineer is configuring a Unity Catalog external location to securely access data in an AWS S3 bucket. The engineer has already created an IAM role with the necessary permissions and configured the storage credential. Which additional step is required to allow Databricks to access the S3 bucket?

Medium
216

A data engineer is setting up Lakehouse Federation to query an external MySQL database from Databricks. The engineer creates a connection using the MySQL connector and a foreign catalog. Users in the 'analysts' group report that they can see the foreign catalog but cannot query any tables. The engineer has granted USAGE on the connection to the 'analysts' group. Which TWO additional permissions must be granted to the 'analysts' group to allow them to query tables in the foreign catalog? (Choose two.)

Hard
217

Refer to the exhibit. A Databricks job fails with a 403 Forbidden error when trying to write to the S3 bucket. Why does this happen?

Hard
218

A data engineer is designing a solution to share a Delta table with an external partner organization. The partner uses a different Databricks account and must be able to read the table, but the data must not be copied outside the provider's cloud storage. The provider uses Unity Catalog and wants to minimize operational overhead while ensuring the partner sees only the shared table. Which Unity Catalog feature should the engineer use?

Hard
219

Which action allows a Data Engineer to receive a Slack notification when a Delta Live Tables pipeline finishes successfully?

Medium
220

A data engineer is managing a Delta Share that includes a table with customer transactions. The share is used by multiple recipients. The engineer needs to update the shared data daily with new transactions and also remove data for customers who have requested deletion (right to be forgotten). The engineer wants to ensure recipients see the updated data without having to recreate the share. What is the best approach?

Hard
221

A data engineer is debugging a slow-running query. They notice that the data is skewed, causing one task to take significantly longer than others. Which approach effectively addresses this skew?

Medium
222

A Data Engineer is developing a Delta Live Tables (DLT) pipeline using Python. They need to ensure that records failing a specific data quality check are dropped, but the pipeline continues to process the remaining valid records. Which expectation syntax should the engineer implement?

Medium
223

When designing a streaming pipeline using Structured Streaming, which THREE of the following are necessary to ensure 'exactly-once' processing semantics in Databricks?

Medium
224

You are migrating a legacy CSV-based ETL process to Databricks. The source CSV files contain inconsistent date formats. Which approach provides the most scalable way to handle these inconsistencies during the bronze-to-silver transformation?

Medium
225

Which of the following is the most secure method for a Data Engineer to provide access to a specific Delta table for a temporary project?

Medium
226

A production Databricks workflow involves a task that runs a notebook. The notebook takes 15 minutes to finish, but the workflow is set to timeout after 10 minutes. What happens?

Medium
227

An organization wants to monitor and limit the spend of their Databricks SQL warehouses. Which feature is most appropriate for setting alerts when costs exceed a certain threshold?

Medium
228

Which THREE of the following are benefits of using Delta Lake over standard Parquet files for your data lake storage?

Hard
229

An organization is migrating to Unity Catalog and needs to secure sensitive data. Which TWO of the following statements regarding Unity Catalog security best practices are correct?

Hard
230

Which TWO of the following are true regarding Unity Catalog's ability to govern external locations?

Hard
231

A data engineer is tasked with ensuring that sensitive information in a 'customer' table is masked for all users except the 'Data_Science' group. What is the correct Unity Catalog feature to implement?

Hard
232

A data engineer wants to monitor the data quality of a Delta table over time. Which tool is most appropriate for this task?

Medium
233

An enterprise data team runs a large nightly batch job using a standard all-purpose cluster. The job frequently fails due to cloud provider spot instance pre-emptions and takes over four hours to complete. How should the engineer refactor this architecture for maximum cost efficiency and reliability?

Medium
234

A Data Engineer is working on a Delta Live Tables (DLT) pipeline that ingests JSON files from cloud storage. The pipeline must drop rows where the 'email' column is null and also flag rows where 'age' is negative as invalid, but still process them. Which combination of DLT expectations should be used?

Medium
235

When configuring a Service Principal to access a Unity Catalog-enabled workspace, which THREE steps are required to ensure secure and functional access?

Hard
236

Which TWO of the following are valid ways to trigger a job in Databricks?

Medium
237

A team has a large Delta table that is rarely updated. What is the most cost-effective way to store this data while maintaining the ability to query it with Databricks SQL?

Easy
238

What is the primary benefit of using a Job Cluster instead of an All-Purpose Cluster for production workloads?

Easy
239

A Data Engineer wants to monitor cluster health proactively. Which metric is most effective for identifying that a cluster needs to be scaled up to handle increasing workload demands?

Medium
240

A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?

Medium
241

A data engineer is optimizing a Spark job that reads from a large Delta table and performs a join with a smaller dimension table. The job is running slowly, and the engineer suspects data skew and shuffle overhead are the main issues. Which two techniques should the engineer apply to improve performance and reduce cost? (Choose two.)

Hard
242

A financial institution is building a Gold layer table that must support point-in-time queries to reconstruct account balances as of any past date. The source data includes transactions with effective dates and an audit log of changes. Which modeling technique is most appropriate?

Medium
243

A data engineering team stores customer transaction data in a Unity Catalog managed table named prod.finance.transactions. The security team requires that any query referencing this table, whether through a view or directly, is recorded with the identity of the user who ran it, and that the audit logs are retained for 365 days. The workspace uses Unity Catalog and has audit logs delivered to a cloud storage location. Which configuration should the data engineer verify or set to meet the requirement that all access to the table is captured with the user identity?

Medium
244

A data engineer is setting up a Delta Share to provide an external partner with access to a subset of data. The partner will use a non-Databricks client that supports the Delta Sharing protocol. Which two actions are required to enable the partner to access the shared data? (Choose two.)

Medium
245

A Spark job is failing with an OutOfMemoryError (OOM) during a group-by operation on a skewed key. What is the most effective way to resolve this?

Medium
246

You are writing a PySpark script to join two large tables. You want to ensure the join operation is optimized for performance by broadcasting the smaller table. Which configuration property should you adjust, or code construct should you use, to force this behavior?

Medium
247

A data engineer is deploying a Databricks Asset Bundle (DAB) that defines a job with a notebook task. The bundle validates locally, but deployment fails with 'Error: cannot find notebook at path /Workspace/Users/dev@example.com/pipeline/ingest'. The engineer confirms the notebook exists in the workspace at that exact path. Which action should the engineer take to resolve the deployment failure?

Hard
248

A Data Engineer is building a Databricks SQL pipeline that ingests clickstream events from a Delta table. The events table contains a nested column `payload` of type STRUCT with fields `page_id` (STRING), `duration` (INT), and `referrer` (STRING). The engineer needs to flatten the `payload` fields into top-level columns and drop any records where `page_id` is NULL. Which SQL expression accomplishes this transformation while preserving all other columns?

Medium
249

A Data Engineer is building a Delta Live Tables (DLT) pipeline to ingest raw JSON data. They need to ensure that records missing the required 'user_id' field are dropped while simultaneously capturing these discarded records in a separate table for auditing purposes. Which approach achieves this in DLT?

Medium
250

A data engineer is designing a Unity Catalog governance model for a new data lakehouse. They need to ensure that data access is auditable and that sensitive data is protected. Which two actions should the engineer take to meet these requirements? (Choose two.)

Medium
251

A data engineer is configuring a Databricks Workflow that must run a notebook task only after a previous task that writes to a Delta table has completed successfully. The engineer wants to ensure that if the first task fails, the second task does not run. Which feature should the engineer use to define this dependency?

Easy
252

A data engineer is setting up a new Unity Catalog metastore. What is the primary purpose of the 'Metastore Admin' role?

Medium
253

An engineer is writing a Python function to process data in a Databricks Notebook. Which command should they use to ensure that secrets, such as API keys, are not hardcoded or exposed in the plain text of the notebook?

Medium
254

A Data Engineer needs to ensure that PII data in a Delta table is accessible only to members of the 'hr_admin' group, while allowing all other users to view the non-PII columns. Which Unity Catalog feature is the most efficient way to implement this requirement?

Medium
255

A Data Engineer is using Lakehouse Federation to query a Snowflake database from Databricks. The engineer has created a connection and a foreign catalog. Which statement correctly describes how data is accessed when a user queries a table in the foreign catalog?

Easy
256

Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?

Hard
257

Refer to the exhibit. You are appending data to an existing Delta table. What is the most likely cause of this error, and how should you resolve it?

Medium
258

A data engineer needs to grant a group of users the ability to run a specific Databricks job but not modify its configuration. The job is managed by a service principal. Which permission level should be assigned to the group on the job?

Easy
259

A Data Engineer is using Lakehouse Federation to query an external PostgreSQL database from Databricks. The engineer creates a foreign catalog named 'pg_catalog' using a connection that specifies the host, port, and credentials. Users report that queries against the foreign catalog fail with a permission error, even though the connection works when tested. The engineer confirms that the connection has the correct credentials and that the PostgreSQL user has SELECT privileges on the required tables. What is the most likely cause?

Hard
260

A Data Engineer is building a Lakeflow Spark Declarative Pipelines pipeline that ingests JSON sensor events. The pipeline must drop records where the `sensor_id` is NULL, ensure that `event_time` is not in the future, and continue processing without failing the update. Which combination of expectations should be used?

Medium
261

A data engineer is troubleshooting a Databricks SQL query that occasionally fails with 'Query exceeded the maximum allowed execution time' on a shared SQL warehouse. The query is a complex aggregation over a large Delta table. The engineer needs to identify the root cause and ensure the query can complete successfully. Which action should the engineer take first?

Hard
262

A data engineer is preparing to deploy a production Databricks Workflow that must be maintainable and auditable. The engineer wants to ensure that changes to the workflow are tracked and that failures can be diagnosed quickly. Which TWO practices should the engineer implement? (Choose two.)

Medium
263

Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?

Hard
264

A data engineer is using Delta Sharing to share a table with an external partner. The partner needs to access the shared data using their own Databricks workspace. Which protocol does Delta Sharing use to enable this cross-platform sharing?

Easy
265

A data engineer is troubleshooting a production Databricks job that intermittently fails with 'SparkOutOfMemoryError'. The job processes large datasets with skewed partitions. The engineer wants to monitor the job to proactively detect memory pressure before failures occur. Which metric should the engineer monitor on the driver and executor nodes?

Hard
266

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?

Hard
267

A data engineer is configuring a Unity Catalog metastore to use a customer-managed key (CMK) for encryption at rest. The engineer has created the necessary Key Vault and key in Azure. Which additional configuration is required to enable CMK for the metastore?

Medium

Frequently asked questions

What does the troubleshooting domain cover on the Databricks-DE-Pro exam?
troubleshooting questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 267 troubleshooting questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only troubleshooting questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.