Courseiva

Databricks-DE-Assoc · domain

scenario questions

Practise Databricks Certified Data Engineer Associate scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

276 questions43 easy162 medium71 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (276)

Click any question to see the full explanation, or start a practice session above.

1

A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?

Medium
2

Which of the following is a fundamental principle of implementing effective CI/CD for Databricks workflows?

Medium
3

A data engineer is tasked with securing sensitive PII data in Unity Catalog. Which THREE actions are recommended to ensure robust security and compliance?

Medium
4

Which of the following describes the correct order of operations to configure a new external location in Unity Catalog?

Hard
5

A data engineer is designing a Delta Live Tables (DLT) pipeline. They need to ensure that records with missing values in the 'customer_id' column are dropped during the ingestion process. Which constraint syntax should be used?

Medium
6

A data engineer is preparing a notebook that must authenticate to cloud storage using a short-lived token that is automatically rotated by Databricks and is never written into the notebook source. The engineer wants the least administrative overhead while keeping secrets out of the code. Which approach should the engineer use?

Medium
7

Your organization requires that all data processing pipelines enforce a strict schema to prevent corrupt data from landing in the Silver layer. Which feature should be configured to ensure that only data matching the expected schema is written?

Medium
8

A CI/CD pipeline uses the Databricks CLI to deploy a job to production. The pipeline must ensure that the job configuration is identical across environments except for the cluster size, which differs between staging and production. Which approach best supports this requirement?

Hard
9

Refer to the exhibit. A security administrator applies this Unity Catalog policy. What is the impact on users in the 'analyst-group'?

Hard
10

Which THREE of the following are core components of the Databricks Intelligence Platform?

Medium
11

A data engineer notices that a production Delta Lake table is experiencing slow read performance during concurrent write operations. The table contains millions of small files. Which action should the engineer take to resolve this performance degradation?

Medium
12

A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)

Hard
13

What is the primary benefit of using Data Live Tables (DLT) for managing dependencies between tables in a pipeline?

Medium
14

A data engineering team uses Databricks Asset Bundles (DABs) to manage a job that must deploy to both a staging and a production workspace. The team wants to avoid hardcoding workspace-specific values such as the cluster ID and the storage path in databricks.yml. Which approach should they use?

Medium
15

Which feature in Unity Catalog is primarily used to track data movement and transformation history for compliance and auditing?

Easy
16

When considering the Databricks Intelligence Platform, what is the primary role of the 'Lakehouse' architecture?

Medium
17

A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?

Medium
18

A data engineer has a Delta table named `sales` with columns `sale_id`, `customer_id`, `amount`, and `sale_date`. They need to create a new table that contains only the `customer_id` and the total `amount` per customer for all sales in 2023. Which SQL statement correctly creates this aggregated table?

Medium
19

A platform team is setting up a CI/CD pipeline that deploys Databricks jobs and notebooks from a Git repository. They must ensure deployments are secure and auditable. (Choose two.)

Hard
20

A data engineering team is migrating a legacy data warehouse to Databricks. They want to ensure that their raw data is ingested into a 'Bronze' table in its original format. Which approach is most recommended for this ingestion layer?

Medium
21

Which TWO of the following scenarios are valid use cases for utilizing Delta Lake's Change Data Feed (CDF)?

Medium
22

A data engineer is configuring a Databricks job that must run on a schedule and send an email notification if the job fails. They want to minimize manual intervention. Which feature should they use to define the schedule and failure notification?

Medium
23

A data engineer needs to run a nightly transformation that reads a large Parquet dataset, writes a curated Delta table, and then immediately runs OPTIMIZE and VACUUM on that table. The engineer wants each step to be observable, retryable, and to avoid data loss if VACUUM fails. Which orchestration approach best meets these requirements?

Hard
24

A data engineer is building a Databricks job that processes millions of small JSON files landed in cloud storage each hour. The job currently spends most of its runtime on file listing and task scheduling overhead. The engineer wants to improve throughput without changing the downstream table schema. Which change should be made to the ingestion step?

Medium
25

A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save costs. Which Trigger should be used in the Structured Streaming code?

Medium
26

Refer to the exhibit. A data engineer is running these commands before performing heavy write operations into a Delta table. What is the primary benefit of enabling these configurations?

Hard
27

A data engineering team needs to restrict access to a sensitive customer table in Unity Catalog so that junior data engineers can only view non-PII columns, while senior engineers can view all columns. Which approach should be used to implement this requirement securely and efficiently?

Medium
28

A data engineer is using Structured Streaming to ingest data from a Kafka topic. They want to ensure that if the pipeline fails, it can resume exactly where it left off, without processing duplicate data. Which component enables this functionality?

Medium
29

Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?

Hard
30

A data engineer is configuring Unity Catalog governance for a multi-department organization. Which TWO actions require the metastore admin or catalog owner to have a workspace-independent metastore assigned? (Choose TWO)

Hard
31

Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?

Medium
32

A data engineer is migrating legacy tables to Unity Catalog. Which TWO of the following statements regarding the transition to three-level namespace (catalog.schema.table) are true?

Hard
33

When ingesting data from cloud storage, what is the most important reason to use a Service Principal or IAM Role instead of individual user credentials?

Medium
34

A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?

Hard
35

A data engineer is configuring a Lakeflow Job that processes sensitive customer data. The job must notify the on-call team when a run fails and must also capture the run's output for auditing. Which TWO actions should the engineer take in the Lakeflow Jobs configuration? (Choose two.)

Medium
36

Which security feature should a data engineer configure to ensure that audit logs from all Databricks workspaces are captured and stored in a single, centralized location?

Medium
37

A data engineer is using Auto Loader to ingest data from a Kafka topic into a Delta table. The engineer wants to ensure that the ingestion handles late-arriving data and provides exactly-once semantics. Which combination of features should the engineer use?

Hard
38

A data engineering team is building a medallion architecture in Databricks. In the Silver layer, streaming data from Kafka must be cleaned, deduplicated, and written into a Delta table. Which Structured Streaming output mode should the engineer select to ensure append-only storage of fully processed, stateful deduplicated records?

Medium
39

Refer to the exhibit. A data engineer is reviewing the configuration of a Delta table that is frequently queried by 'customer_id'. Given the current metadata, which action will provide the most significant improvement to query performance?

Medium
40

A data engineer needs to store structured data in a cloud object storage location while maintaining full ACID guarantees. Which storage format is the foundation of the Databricks Lakehouse architecture that enables this functionality?

Easy
41

A data engineer has a Lakeflow Job with two tasks: Task1 and Task2. Task2 must run only if Task1 succeeds. The engineer also wants Task2 to be skipped if Task1 fails, but the overall job status should be marked as failed. Which configuration should the engineer use for the dependency between Task1 and Task2?

Hard
42

A data engineer is configuring a Databricks cluster to run a Spark job that processes large datasets. The job requires high memory and will run for several hours. The engineer wants to minimize costs while ensuring the job completes successfully. Which cluster configuration should the engineer choose?

Medium
43

A data engineer needs to pass the execution date to a job task dynamically. Which feature should they use?

Medium
44

Refer to the exhibit. A data engineer attempts to run the GRANT statement above in a Databricks SQL query editor. What is the most likely reason for failure?

Medium
45

A team is setting up a CI/CD pipeline that deploys Databricks Asset Bundles to production. They want the pipeline to be secure and to fail fast before any production resources are changed. Which TWO practices should they implement? (Choose two.)

Medium
46

A data engineer is working with a large, partitioned table and needs to perform a complex transformation. Which THREE of the following strategies will optimize query performance for this transformation?

Hard
47

When promoting code from a development workspace to a production workspace, what is the primary risk of using manual notebook exports?

Medium
48

A data engineer needs to create a Silver Delta table that contains only distinct, non-null `customer_id` values from a Bronze table, and the result must be refreshed idempotently each night. Which statement best satisfies the requirement?

Easy
49

Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?

Medium
50

A data engineer is working on a Bronze-to-Silver transformation. Which THREE of the following practices are recommended to optimize performance and data quality during this stage?

Medium
51

Refer to the exhibit. A user who is a member of 'data_analysts_group' reports they cannot see the 'orders' table in the Catalog Explorer. What is the most likely cause?

Medium
52

What is the primary purpose of a 'feature branch' in a Git-based workflow for Databricks?

Easy
53

A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?

Medium
54

A data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?

Medium
55

A data engineering team is setting up a CI/CD pipeline for Databricks notebooks using Databricks Repos and a Git provider. They want to ensure that changes are tested before being merged and that production deployments are controlled. Which TWO practices should they implement? (Choose two.)

Hard
56

A data engineer is creating a Lakeflow Job that must run a Python script stored in DBFS. The engineer wants to ensure the script is executed with the correct dependencies and environment. Which task type should be used?

Easy
57

A data engineer needs to share a subset of a table with an external partner who does not have access to the internal Databricks workspace. What is the most appropriate method to achieve this?

Medium
58

A data engineer is working with a Delta table that contains a column 'timestamp' of type timestamp. The table is partitioned by date. The engineer needs to run a query that filters on a specific date range and also on a high-cardinality column 'user_id'. The query is performing poorly. Which optimization technique should the engineer apply to improve query performance?

Hard
59

A data engineer is designing a secure architecture using the Databricks Intelligence Platform. Which TWO of the following statements accurately describe the role and capabilities of Unity Catalog within this platform? (Choose TWO)

Hard
60

A data engineer is building a Lakeflow Job that must process a parameterized date range. The engineer wants to pass start_date and end_date values into a notebook task at runtime and have those values available as widget-like parameters inside the notebook. Which approach should the engineer use?

Medium
61

A data engineer is configuring audit logging for a Databricks workspace that uses Unity Catalog. The security team requires that all access to data in Unity Catalog be logged and available for analysis in a centralized location. The data engineer wants to enable the delivery of audit logs to an AWS S3 bucket. Which configuration should the data engineer use?

Medium
62

A data engineer is using Auto Loader to stream data from Kafka into a Delta table. The Kafka topic receives messages in Avro format, and the schema is stored in a Confluent Schema Registry. The engineer wants Auto Loader to automatically fetch the schema from the registry and evolve it as new versions are registered. Which configuration should the engineer use?

Hard
63

Refer to the exhibit. A job fails with the provided error message. What is the most likely cause of this failure in a Databricks Delta Lake environment?

Hard
64

An organization is adopting the Databricks Intelligence Platform and wants to leverage Mosaic AI for building custom machine learning models. Which feature allows data engineers to track machine learning experiments, log parameters, and manage model artifacts reliably?

Medium
65

A data engineer is configuring a Databricks Auto Loader stream to ingest CSV files from a cloud storage location into a Delta table. The CSV files have a header row, and the engineer wants to automatically infer the schema and store the inferred schema in a specified location for consistency across restarts. Which Auto Loader option should be used to persist the inferred schema?

Medium
66

A data engineer is managing a Unity Catalog table that contains sensitive financial data. The table is owned by the 'finance' group, and the data engineer needs to allow the 'auditors' group to read the table but not modify it. Additionally, the data engineer wants to ensure that the 'auditors' group can see the table's metadata (e.g., column names and types) but cannot access the underlying data files directly. Which TWO actions should the data engineer take to meet these requirements? (Choose two.)

Medium
67

A data engineer needs to allow a service principal to read data from a Unity Catalog table `sales.orders` and also write to a volume `sales.raw_data`. Which set of privileges should be granted to the service principal?

Medium
68

A data engineering team runs a nightly Lakeflow Job that ingests files from cloud storage, transforms them with a notebook, and then runs a SQL task. The team wants the SQL task to execute only after the notebook transform succeeds, but they do not want the SQL task to wait for a fixed delay. Which Lakeflow Jobs feature should they configure on the SQL task?

Medium
69

A data engineer needs to grant the `analyst` group the ability to read data from a Unity Catalog table `sales.fact_orders`. They also want to ensure that members of `analyst` can see the table in the catalog explorer but cannot modify it. Which privilege should be granted to the `analyst` group on the table?

Easy
70

A team is designing a CI/CD pipeline using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for managing Databricks projects?

Medium
71

A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?

Medium
72

A data engineer is designing a Bronze-to-Silver transformation pipeline using Delta Lake. They need to ensure that the Silver table contains only records where the 'transaction_id' is not null and the 'amount' is positive. Which technique best ensures data quality at this stage?

Medium
73

A job is failing with a 'Disk Space' error on the worker nodes. The code performs several large joins. Which configuration should the engineer adjust to mitigate the disk space usage?

Medium
74

When implementing a Medallion Architecture, what is the primary purpose of the 'Silver' layer?

Easy
75

A data engineer is using Delta Live Tables to build a pipeline. They need to create a table that contains the latest record for each customer based on a `last_updated` timestamp. The source is a streaming table with append-only data. Which Delta Live Tables operation should be used to achieve this?

Medium
76

Which THREE of the following are benefits of using Delta Live Tables (DLT) for managing your data pipelines?

Hard
77

A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?

Hard
78

A data engineer runs a nightly Databricks job that ingests data into a Delta table using multiple concurrent write streams. The engineer notices that some write transactions are failing with a ConcurrentAppendException. The job writes to the same partition of the table from different tasks. Which action should the engineer take to resolve this issue?

Medium
79

A data engineer is troubleshooting a slow-running Spark job on Databricks. Which TWO metrics in the Spark UI are most useful for identifying data skew?

Medium
80

Refer to the exhibit. A CI/CD pipeline running a Databricks CLI command fails with the error shown. What is the most likely cause?

Medium
81

A data engineer is designing an access control model in Unity Catalog for a new catalog `finance`. The team wants to follow the principle of least privilege while still enabling collaboration. Which TWO of the following practices best align with Unity Catalog's privilege model? (Choose two.)

Medium
82

What is the primary function of a Delta Lake 'Vacuum' operation?

Easy
83

A data engineer is monitoring a Databricks job and notices that the job's duration has gradually increased over the past week. The job reads a large Delta table, performs aggregations, and writes results to another Delta table. The engineer wants to identify the stage that is taking the most time. Which Spark UI tab should the engineer use to quickly identify the slowest stage?

Easy
84

A data engineer is investigating why a Databricks job that writes to a Delta table is experiencing performance degradation over time. The job performs frequent small appends. Which TWO actions should the engineer take to improve write performance? (Choose two.)

Hard
85

A data engineer needs to share a table in Unity Catalog with a partner organization using a different Databricks account. Which feature should be used to provide secure access without moving the data?

Medium
86

An engineer needs to identify the root cause of a job failure. Which THREE of the following are valid locations or methods to investigate the logs?

Medium
87

A data platform team is migrating their deployment process to Databricks Asset Bundles (DABs). They already have a Python wheel task defined in a Databricks Job and a set of notebooks in a Git repository. They want the bundle deployment to be repeatable across development, staging, and production targets with different cluster sizes. Which approach should they take to parameterize the target-specific cluster configuration?

Medium
88

A data engineer wants to use 'Expectations' in Delta Live Tables to monitor data quality. What happens if a record violates an expectation defined with the 'fail' constraint?

Medium
89

A data engineer has a Lakeflow Job with three tasks: bronze_ingest, silver_transform, and gold_aggregate. The silver_transform task must run only if bronze_ingest succeeds, and gold_aggregate must run only if silver_transform succeeds. The engineer also wants gold_aggregate to run even if silver_transform fails, so that partial results can be published. Which configuration should the engineer apply to gold_aggregate?

Hard
90

A data engineer is working on a Delta Lake table that has accumulated millions of small files due to frequent streaming updates. This fragmentation has significantly degraded query performance. Which operation should the engineer execute to optimize file layout without altering table data?

Medium
91

Which of the following is the primary benefit of using a 'Job Cluster' rather than an 'All-Purpose Cluster' for running scheduled data pipelines?

Easy
92

A data engineer has a Unity Catalog table `prod.sales.orders` that contains a column `customer_email`. They need to allow analysts in the `marketing` group to query the table but only see a masked version of `customer_email` (e.g., `a***@example.com`). The masking logic is implemented as a SQL user-defined function `prod.security.mask_email`. Which statement should the engineer execute to apply the mask?

Medium
93

A data engineering team stores all production notebooks and job definitions in a Git repository. They want every merge to the `main` branch to automatically deploy the updated notebooks to the production Databricks workspace without any manual copy/paste. Which approach should they implement?

Medium
94

A data engineer needs to troubleshoot a job that is failing during the 'shuffle' phase. Which Spark UI tab should the engineer examine to analyze the shuffle partitions and identify potential imbalances?

Medium
95

A data engineer maintains a Delta table named inventory.products with columns product_id, category, price, and updated_at. The engineer needs to create a new table that contains one row per category with the average price and the most recently updated product_id in that category. The query must be efficient and use only standard Databricks SQL. Which statement should the engineer run?

Medium
96

A data engineer is creating a Silver table in a Delta Live Tables pipeline. The pipeline must continuously ingest new files from a cloud storage location as they arrive, and the engineer wants to avoid reprocessing files that were already ingested. Which approach should be used to read the source data?

Easy
97

An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?

Medium
98

A data engineer is building a Gold aggregate table that summarizes daily sales by product category. The Silver source is a streaming Delta table that receives late-arriving events up to 48 hours old. The engineer needs the Gold table to always reflect the most accurate aggregates, including corrections for late data, without full recomputation. Which approach is most appropriate?

Hard
99

Refer to the exhibit. An engineer notices that queries filtering by 'customer_id' are running slowly on the 'orders' table. Based on the exhibit, what is the most appropriate action to resolve this?

Medium
100

A data engineer is using Databricks Jobs to run a nightly ETL pipeline. The job occasionally fails due to a transient network error when writing to an external database. The engineer wants to automatically retry the job a few times before marking it as failed. What is the most efficient way to configure this in Databricks?

Medium
101

Refer to the exhibit. If 'task1' fails due to a timeout, what happens to 'task2'?

Hard
102

Which component of Databricks CI/CD is responsible for executing automated tests on code before it is merged into the main branch?

Easy
103

Which command is used to query the history of a Delta table to perform time travel?

Medium
104

A data engineer manages a Unity Catalog table `sales.raw.transactions` that contains a column `credit_card_number`. The security team requires that users in the `auditors` group can see the full credit card number, while all other users who have SELECT privileges on the table should see only the last four digits. The engineer decides to use a column mask. Which SQL statement should the engineer execute to meet this requirement?

Hard
105

A data engineer is designing a Delta Lake Bronze-to-Silver pipeline in Databricks and needs to ensure that downstream consumers receive high-quality data. Which TWO data quality enforcement mechanisms are natively supported in Delta Live Tables using expectations?

Hard
106

A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?

Easy
107

Which of the following is the best practice for managing service principals in a Databricks workspace?

Medium
108

Refer to the exhibit. A data engineer is configuring a streaming pipeline. Which outcome will these specific configurations have on the target table's performance?

Medium
109

A team is implementing a CI/CD process for their Delta Live Tables (DLT) pipelines. Which THREE of the following practices are recommended to ensure reliable deployment?

Hard
110

Refer to the exhibit. A CI/CD pipeline fails with the provided error. What is the most likely cause?

Hard
111

A team wants to automate deployment of Databricks jobs from a Git repository using a CI/CD pipeline. They need a tool that reads a declarative project definition and creates or updates jobs, pipelines, and notebooks in a target workspace. Which Databricks capability should they use?

Easy
112

A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?

Hard
113

An engineer notices that a specific notebook job is consistently taking longer to start. They observe high 'initialization' times in the job logs. Which action should the engineer take to improve startup time?

Medium
114

A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)

Hard
115

A data engineer is monitoring a Databricks job and notices that the job's tasks are spending a significant amount of time in garbage collection (GC). The job processes large amounts of data with many small objects. Which action should the engineer take to reduce GC overhead?

Medium
116

A data engineer is using Databricks SQL to analyze data stored in a Delta table. The engineer wants to optimize query performance by leveraging Delta Lake features. Which TWO actions should the engineer take to improve query performance on the Delta table? (Choose two.)

Medium
117

A data engineer runs a Structured Streaming job that writes to a Delta table. The job processes data from a Kafka topic and uses a 10-minute watermark. After a few hours, the engineer notices that the streaming query's input rate is steady, but the processing rate has dropped significantly, and the batch duration has increased from 5 seconds to over 2 minutes. The job is running on a cluster with autoscaling enabled. Which action should the engineer take FIRST to diagnose the performance degradation?

Medium
118

A data engineer needs to share a table in Unity Catalog with external partners who do not have access to the Databricks workspace. Which feature should the data engineer configure?

Medium
119

Which TWO of the following statements are correct regarding the use of Databricks Repos for CI/CD?

Medium
120

Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?

Easy
121

A data engineer is working with a Delta table that contains a column named raw_data of type STRING, which holds JSON strings. The engineer needs to extract specific fields from this JSON and store them as separate columns in a new Delta table. Which approach is most efficient and maintains data quality?

Easy
122

A data engineer is optimizing a Delta table that suffers from slow read performance due to small file sizes. Which command should the engineer execute to consolidate these small files into larger, more efficient files without altering the underlying table data?

Medium
123

A data engineer needs to monitor the costs associated with specific projects running on a shared Databricks workspace. Which feature should the engineer use to attribute these costs accurately?

Easy
124

You are tasked with handling late-arriving data in a streaming pipeline that performs windowed aggregations. Which approach ensures that the output remains accurate while balancing memory usage?

Medium
125

A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?

Medium
126

An analytics team needs to frequently query a large Delta table by a high-cardinality customer_id column and a date column. To optimize query performance and reduce data scanning during filtering, how should the data engineer structure the table layout?

Medium
127

A data engineer has a Unity Catalog table `prod.sales.orders` containing a column `customer_email`. Company policy requires that users in the `analyst_group` see only the domain part of the email (e.g., `***@example.com`), while members of `pii_admin_group` must see the full email. The engineer wants to enforce this at query time without creating separate views. Which approach should the engineer use?

Medium
128

A data engineering team stores its Databricks notebooks and Python files in a Git repository. They want to avoid manually copying files into the workspace and ensure that the production workspace always runs the exact code version that passed tests. Which approach should they use?

Medium
129

A data engineer is optimizing a Databricks job that reads from a large Delta table and performs a join with a smaller table. The job is experiencing performance issues due to shuffling. The engineer wants to reduce the amount of data shuffled during the join. Which technique should the engineer use?

Hard
130

A data engineering team stores notebooks in a Git repository and wants automated deployments to a Databricks workspace. They configure a GitHub Actions workflow that runs a Databricks CLI command to deploy Databricks Asset Bundles. The workflow authenticates using a service principal OAuth token stored in GitHub Secrets. After the first successful run, subsequent runs fail with an authentication error. The token was created with a 1-hour lifetime. What should the team do to ensure the workflow can authenticate reliably on every run?

Medium
131

A data engineer is configuring a storage credential in Unity Catalog to access an AWS S3 bucket. The engineer creates an IAM role with the necessary permissions and sets up the storage credential using the role ARN. What additional step is required to allow Databricks to assume the role?

Medium
132

A data engineer stores a Delta table in Unity Catalog at the managed location of the schema 'sales'. The table is later dropped using DROP TABLE. What happens to the underlying data files?

Medium
133

A data engineer is creating a Lakeflow Job that must run a notebook every weekday at 06:00 in the company's local time zone, which is America/New_York. The engineer configures a schedule trigger but the job runs at the wrong time. Which setting should the engineer verify first?

Easy
134

A team is preparing to optimize their Databricks data transformation pipeline. Which THREE of the following actions are considered best practices for optimizing Delta Lake performance?

Medium
135

A team is using Databricks Asset Bundles (DABs) to manage their CI/CD pipeline. They want to run unit tests on their Python code before deploying the bundle. Where should the tests be executed in the pipeline?

Medium
136

A data engineer has a Lakeflow Job that runs daily. They want to receive an email only when the job fails, not on every run. Which notification configuration should they set?

Easy
137

A data engineer is designing a pipeline on Databricks that requires ACID transactions and schema enforcement for streaming data. Which storage abstraction should they use to ensure data reliability and support time travel?

Medium
138

A data engineer wants to pass a file path from Task A to Task B in a Lakeflow Job. Task A is a notebook that computes the path, and Task B is a notebook that reads from that path. Which mechanism should the engineer use to share the value between tasks?

Medium
139

Which Databricks feature should be used to securely share data with external organizations without duplicating the data?

Medium
140

A data engineer has a Lakeflow Job with a linear dependency chain: Task A, then Task B, then Task C. Task B sometimes fails due to transient errors. The engineer wants Task C to run only if Task B succeeds, but also wants Task B to be retried automatically before considering the job failed. Which configuration should they use?

Hard
141

Which object type in Databricks Unity Catalog acts as the top-level container for organizing schemas and tables, providing a unified namespace for data assets?

Easy
142

Why is it important to use Service Principals instead of Personal Access Tokens (PATs) for CI/CD automation in Databricks?

Easy
143

Which of the following is a recommended strategy for managing library dependencies in a CI/CD pipeline for Databricks?

Medium
144

A data engineer needs to grant a group `data_consumers` the ability to query a view `prod.reporting.sales_summary` that is defined on top of tables in the same catalog. The group currently has no privileges on the underlying tables. What is the minimum set of privileges the engineer must grant to `data_consumers` so they can query the view successfully?

Medium
145

A data engineer is analyzing a Spark job that is failing with 'Out of Memory' (OOM) errors. Which configuration parameter should be tuned to increase the amount of memory allocated to the execution of joins and aggregations?

Medium
146

Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?

Easy
147

Which of the following is the primary benefit of using 'Auto Loader' (cloudFiles) for data ingestion in Databricks compared to standard batch processing?

Easy
148

Refer to the exhibit. When is this job scheduled to run?

Hard
149

An organization requires that all data access logs across multiple Databricks workspaces be captured and sent to a centralized security information and event management (SIEM) system. Which Databricks feature should be configured to capture these audit events?

Medium
150

A data engineer is configuring a Unity Catalog storage credential to access an AWS S3 bucket. The S3 bucket policy grants access to an IAM role. The data engineer creates a storage credential with that IAM role's ARN. However, when attempting to create an external location using this storage credential, the operation fails with an error indicating insufficient permissions. The data engineer verifies that the IAM role has the correct S3 permissions. What is the most likely cause of the failure?

Hard
151

A data engineer is building a Silver table in Delta Lake from a Bronze table that contains raw JSON events. The engineer needs to flatten a nested struct column named 'device' with fields 'type' and 'os', and also extract a field from an array of structs named 'events'. The goal is to produce a clean, denormalized Silver table. Which PySpark operation should the engineer use to achieve this transformation efficiently?

Medium
152

A data engineering team is implementing a CI/CD pipeline for Databricks notebooks and jobs using Databricks Asset Bundles. They want to ensure deployments are reproducible and that production changes are traceable. Which TWO of the following practices should they follow? (Choose two.)

Hard
153

Refer to the exhibit. A data engineer is using this command to reload data into a Delta table after a schema correction. What is the primary effect of setting the 'force' option to 'true' in this context?

Medium
154

A data engineer is setting up a Lakeflow Job that runs a notebook task. The engineer needs the task to always execute even if the upstream task in the workflow fails. Which configuration should be applied to the dependent task's condition?

Medium
155

A data engineer is designing a pipeline and needs to ensure that data remains consistent during concurrent read and write operations. Which Databricks feature provides the mechanism to track and validate these operations?

Medium
156

A data engineering team must implement dynamic row-level filtering on a customer analytics table in Unity Catalog so that regional analysts only view records corresponding to their assigned territory. Which TWO steps are required to achieve this using Unity Catalog features? (Choose 2)

Hard
157

What is the primary goal of implementing 'environment parity' in a Databricks CI/CD pipeline?

Medium
158

A data engineer wants to monitor the health and performance of Databricks Jobs over time. Which feature should they use to visualize trends, such as job success rates and average execution times, across multiple runs?

Medium
159

A data engineering team needs to ingest streaming data from Kafka into a Delta table while maintaining exactly-once processing guarantees and low latency. Which Databricks Intelligence Platform feature should they utilize to build this streaming pipeline declaratively?

Medium
160

Refer to the exhibit. The streaming pipeline is experiencing memory issues because the state store size increases continuously. What is the most effective way to address this while maintaining aggregation accuracy?

Hard
161

A data engineer needs to ensure that all queries against a Unity Catalog table are logged for audit purposes. The logs must include the user identity, the query text, and the timestamp. Which Databricks feature should be enabled to capture this information?

Easy
162

A data engineer is migrating legacy batch jobs to Delta Live Tables (DLT). They want to optimize performance for a complex join operation between two large tables. Which TWO strategies should they implement to improve the join efficiency?

Medium
163

A team is using the Databricks CLI in their CI/CD pipeline to deploy jobs and notebooks. They need the pipeline to authenticate to a production workspace without embedding a personal user's credentials. Which authentication method should they configure for the CLI?

Easy
164

Which TWO of the following statements accurately describe the functionality of Unity Catalog within the Databricks Intelligence Platform?

Medium
165

A data engineer has configured a storage credential in Unity Catalog to access an AWS S3 bucket. The credential uses an IAM role with a trust policy that allows Databricks to assume it. The engineer now needs to ensure that only a specific set of users can create external tables pointing to that S3 bucket. Which Unity Catalog object should be used to control this access?

Hard
166

When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?

Easy
167

A data engineer is investigating a Databricks job that failed overnight. The job's status in the Jobs UI shows 'Failed', and the engineer needs to view the error message and stack trace to determine the cause. Where should the engineer look to find the detailed error information for the failed run?

Easy
168

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data pipeline development?

Medium
169

A data engineering team wants to implement Git integration for their Databricks notebooks. Which workflow is considered the best practice for CI/CD in Databricks Repos?

Medium
170

A data engineer is setting up a CI/CD pipeline that runs unit tests on transformation logic before deploying notebooks to a production Databricks workspace. The tests must run quickly and not depend on a live Databricks cluster or external data sources. Which approach best meets these requirements?

Medium
171

A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?

Medium
172

Which THREE of the following represent core security principles enforced by Unity Catalog?

Medium
173

You are building a pipeline where a Bronze table contains JSON data with a nested 'user_info' struct. You need to promote this to a Silver table where 'user_id' is a top-level column. Which approach is the most efficient for this transformation?

Medium
174

A data engineer is building a Delta Live Tables pipeline that ingests streaming data from a Kafka topic into a bronze table, then applies a series of transformations to produce a silver table. The engineer notices that the pipeline is reprocessing all data from the beginning of the Kafka topic on each run, causing high latency. The Kafka topic has a retention period of 7 days, and the pipeline is configured to use the default settings. What is the most likely cause of this behavior?

Hard
175

A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?

Medium
176

A data engineer is asked to ensure that all queries against a Unity Catalog table are recorded for auditing, including the identity of the user and the query text. Which Databricks feature should the engineer enable to capture this information?

Easy
177

Which THREE of the following are primary components of the Databricks Lakehouse architecture?

Medium
178

Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?

Medium
179

A data engineering manager wants to ensure that all production code in Databricks is fully audited and versioned. Which TWO of the following steps are required?

Hard
180

A data engineer is using Databricks Auto Loader to ingest JSON files from a directory. The stream is configured with `cloudFiles.schemaLocation` set to a specific path. The engineer notices that the stream fails when a new file contains an additional column. What is the most likely reason for the failure?

Hard
181

A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?

Medium
182

A data engineer configures a Lakeflow Job to run a notebook task on a job cluster. The notebook reads a parameter named run_date using the widget API. During a manual run, the engineer wants to supply a specific date without editing the notebook. The job also runs on a nightly schedule where the date should default to the current day. Which approach correctly supplies the parameter for both the manual and scheduled runs?

Hard
183

A data engineer needs to share a Delta table managed by Unity Catalog with external partners who do not have access to the Databricks workspace. Which feature should be used to securely grant read-only access to this table without replicating the data?

Medium
184

A data engineer is designing a Delta Lake pipeline that processes streaming sales transactions. The schema evolves frequently, and the pipeline must handle these changes without manual intervention. Which feature should the engineer enable to support automatic schema updates while preventing data corruption?

Medium
185

What is the primary role of a 'Metastore Admin' in a Databricks Unity Catalog environment?

Medium
186

A team uses Databricks Asset Bundles to define a job that must exist in both a staging and a production workspace with different cluster sizes. They want a single bundle definition that deploys correctly to both targets. Which configuration should they use?

Medium
187

A data engineer is setting up a Databricks job that runs a notebook on a schedule. The job must process data from a source that is updated daily and write results to a Delta table. The engineer wants to ensure that if the job fails, it automatically retries up to three times. Which feature should the engineer configure in the job settings to achieve this?

Medium
188

A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?

Medium
189

Your team is migrating a manual Databricks job to a CI/CD pipeline. You need to ensure the job configuration is version-controlled and deployed programmatically. Which approach aligns with Databricks best practices?

Medium
190

Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?

Medium
191

A data engineer is designing an ingestion pipeline using Databricks Auto Loader to process JSON files from an S3 bucket. The pipeline must handle schema evolution and ensure that data is ingested exactly once. Which two features of Auto Loader support these requirements? (Choose two.)

Hard
192

A data engineer wants to move data from a 'Bronze' table to a 'Silver' table while performing data cleaning. They want to ensure that this process is only executed once per data batch. Which approach is best for this requirement?

Medium
193

Refer to the exhibit. A data engineer receives this error when collecting data from a large transformation back to the driver node. Which approach should be used to fix this issue?

Hard
194

Which TWO of the following statements accurately describe the behavior of Delta Lake table constraints and enforcement?

Hard
195

A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing mode to file notification mode. What is the primary reason for making this change?

Hard
196

What must be configured to allow Databricks to access cloud storage on behalf of a user without the user needing to provide their own cloud credentials?

Medium
197

A data engineer needs to join two massive datasets. One dataset is very small (10MB), and the other is very large (1TB). To ensure the join operation is performed as efficiently as possible, which join strategy should be enforced?

Medium
198

Refer to the exhibit. A Databricks administrator is using the Unity Catalog JSON policy to manage access. If the 'data_scientist_1' user attempts to execute an 'UPDATE' command on the 'orders' table, what will be the result?

Hard
199

A data engineer has a Lakeflow Job with a notebook task that occasionally fails due to transient network errors when reading from an external REST API. The engineer wants the task to automatically retry up to three times, but only for this specific task, without affecting other tasks in the job. What should the engineer do?

Medium
200

A data engineer is working in a Databricks workspace where Unity Catalog is enabled. They need to run a SQL query that reads from the table sales in the catalog prod and schema marketing. Which fully qualified name should they use?

Easy
201

Which capability is provided by Databricks' integration with MLflow?

Easy
202

While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?

Medium
203

A data engineer is using Databricks Asset Bundles to deploy a data pipeline that includes a job and a notebook. The engineer wants to ensure that the deployment is idempotent and can be rolled back if needed. Which TWO of the following statements accurately describe the benefits of using Databricks Asset Bundles for this scenario? (Choose two.)

Medium
204

Refer to the exhibit. An organization needs to ensure that members of the 'analyst_group' can only view data where the region column equals 'US'. Based on the provided configuration, what is the best approach to implement this in Unity Catalog?

Medium
205

Refer to the exhibit. An engineer applies these configurations to a cluster. What is the primary benefit of enabling the Databricks IO Cache for a workload that involves repeatedly reading the same Delta tables?

Hard
206

Refer to the exhibit. A data engineer is attempting to run a VACUUM command on a table, but the command fails with an error indicating that the retention period is too short. Given the configuration, what is the most appropriate action the engineer should take to safely remove files older than 7 days?

Hard
207

A data engineer is optimizing a Databricks job that processes a large dataset. The job performs a join between a large Delta table and a small dimension table, then writes the result to a Delta table. The engineer notices that the join is causing a large shuffle and wants to reduce shuffle overhead. Which two actions should the engineer take to improve performance? (Choose two.)

Medium
208

A data engineer is building a Lakeflow Job that must run a sequence of tasks across different compute types. The ingest task must run on a job cluster with a specific Spark configuration, the transform task must run as a Delta Live Tables pipeline, and the report task must run on a separate SQL warehouse. Which TWO statements about task-level compute configuration in Lakeflow Jobs are correct? (Choose two.)

Hard
209

A team is using Databricks Asset Bundles (DABs) to deploy a job to multiple environments (dev, staging, prod). They need to ensure that the job uses different cluster sizes and schedules per environment. Which DABs feature should they use?

Hard
210

A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?

Hard
211

A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.json. The engineer wants to automatically add a column 'date' to the ingested data based on the file path. Which Auto Loader feature should be used?

Medium
212

A data engineer is building a Lakeflow Job with a task that runs a SQL notebook. The task must run only on weekdays and must be completed before 9 AM. The engineer wants to configure the schedule to meet these requirements. What should the engineer do?

Medium
213

A data engineer needs to optimize the layout of a massive Delta Lake table that suffers from poor query performance due to a large number of small files and unsorted data records. Which TWO operations should the engineer execute to resolve these performance bottlenecks?

Medium
214

A data engineer is designing a pipeline on Databricks to process streaming data. Which architectural component acts as the unified storage layer, allowing both batch and streaming workloads to access the same underlying data files in a data lake?

Medium
215

Which feature of the Databricks Intelligence Platform allows users to manage fine-grained access control across workspaces for tables, files, and machine learning models?

Easy
216

A data engineer is troubleshooting a Databricks job that intermittently fails with a `SparkException: Job aborted due to stage failure: Task not serializable`. The job reads from a Parquet file, performs a transformation using a custom function defined in a Python class, and writes to a Delta table. The engineer suspects that the custom function is causing the issue. Which action should the engineer take to resolve the serialization error?

Hard
217

Refer to the exhibit. A data engineer encounters this error while trying to list tables in a schema. Based on the error, what must the engineer do to resolve it?

Hard
218

Which THREE of the following are benefits of using Unity Catalog over the legacy Hive Metastore?

Medium
219

A data engineer has a Unity Catalog table named `sales.raw.orders` that contains a column `credit_card` with sensitive data. The security team requires that users in the `analyst` group see only the last four digits of the credit card number when querying this table, while all other users with appropriate privileges see the full value. The data engineer wants to implement this with minimal disruption to existing queries. Which approach should the data engineer take?

Medium
220

A data engineer is investigating why a Databricks job that reads from a Delta table is slow. The job performs a simple SELECT with a filter on a partition column. The engineer suspects that the table has many small files. Which Spark UI tab should be examined to confirm the number of files read?

Easy
221

A data engineer is configuring a CI/CD pipeline for Databricks notebooks using GitHub Actions. They need to authenticate to Databricks to deploy notebooks. Which authentication method is recommended for production CI/CD pipelines?

Medium
222

What is the primary function of the 'Retries' setting in a Databricks Job task?

Medium
223

A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?

Medium
224

A team stores Databricks notebooks and job definitions in a Git repository and wants every merge to the main branch to automatically deploy to production. Which combination of practices should the pipeline implement to achieve this safely?

Medium
225

Which TWO of the following statements correctly describe the behavior of the Delta Lake 'MERGE' operation when handling schema evolution?

Hard
226

When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?

Medium
227

A data engineer is designing a workflow that requires running a series of tasks with dependencies. The workflow must be able to retry failed tasks automatically and send notifications on failure. The engineer also needs to ensure that the workflow can be triggered on a schedule and via an API. Which Databricks feature should the engineer use?

Hard
228

A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?

Hard
229

A data engineer is working with a large Delta table and notices that queries filtering by 'region_id' are performing slowly. The table is currently partitioned by 'date'. Which strategy should the engineer use to optimize query performance for 'region_id' filtering without increasing the number of partitions?

Medium
230

A data engineer needs to configure a Databricks Job to orchestrate a data pipeline that includes a Python task, a SQL task, and a notebook task. The pipeline requires passing a dynamic run identifier from the Python task to the subsequent SQL and notebook tasks. Which mechanism should the data engineer use to achieve this task-to-task dependency parameter passing?

Medium
231

Refer to the exhibit. The merge operation is failing in your pipeline. What is the root cause of this error, and how should it be resolved?

Hard
232

When utilizing Databricks Asset Bundles, how should secrets (e.g., API keys, database credentials) be handled to ensure security during CI/CD?

Hard
233

A data engineer is configuring Unity Catalog row-level security on a table `prod.finance.transactions`. They want to ensure that users in the `us_team` group can only see rows where `region = 'US'`, while users in the `eu_team` group can only see rows where `region = 'EU'`. They create a row filter function `prod.security.region_filter`. Which TWO statements accurately describe how to implement and manage this row-level security? (Choose two.)

Hard
234

Refer to the exhibit. A data engineer is configuring an IAM policy for a Databricks storage credential. Which of the following is the most significant security risk in this configuration?

Hard
235

Why should developers avoid using 'notebook' references in production pipelines that point to the 'Shared' folder for shared development work?

Medium
236

A data engineer working in a Databricks workspace needs to create a new notebook that will be shared with teammates in the same workspace. They want the notebook to be organized in a folder named 'team_project' and be visible to all workspace users. Which action should the data engineer take to accomplish this in the Databricks workspace?

Easy
237

When integrating Databricks with a CI/CD tool like GitHub Actions, how should you securely manage the authentication token used for deployments?

Easy
238

Which of the following best describes the purpose of 'Credential Passthrough' in Databricks?

Hard
239

A data engineer is using Auto Loader to ingest JSON files from a cloud storage directory into a Delta table. The directory receives files continuously, and the engineer wants to ensure that the ingestion process can handle schema drift where new columns are added to the JSON files over time. The engineer also wants to minimize the need to reprocess all data when the schema changes. Which configuration should the engineer use?

Hard
240

A data engineer manages a Lakeflow Job that runs a long-running notebook task on a job cluster. The task occasionally fails due to transient cloud storage errors, and the engineer wants the task to retry automatically without failing the entire job on the first attempt. The engineer also wants to be alerted only if all retries are exhausted. Which configuration should the engineer apply?

Medium
241

A data engineer is implementing a Type 2 slowly changing dimension in Delta Lake for a customers table. The table has columns customer_id, name, address, effective_date, end_date, and is_current. When a customer's address changes, the engineer wants to expire the existing current row and insert a new current row in a single atomic operation. Which Delta Lake feature should the engineer use?

Medium
242

Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?

Medium
243

A team is transitioning to the Databricks Intelligence Platform. Which TWO actions are required to successfully register a table in Unity Catalog using the three-level namespace?

Hard
244

A CI/CD pipeline must run unit tests on Python transformation code before deploying a Databricks job. The tests should execute quickly without starting a cluster and should validate the transformation logic in isolation. Which approach best meets these requirements?

Hard
245

A data engineer needs to share a Delta table with an external partner who does not have a Databricks account. The engineer wants to provide read-only access to the table for a limited time, ensuring the partner cannot access any other data. Which Databricks feature should the engineer use?

Medium
246

A data engineer maintains a Delta table where each row represents a customer record, and updates arrive continuously as change data capture events. The engineer needs to apply inserts, updates, and deletes from a staging table into the target table in a single atomic operation, matching records on customer_id. Which Delta Lake operation should be used?

Hard
247

A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?

Medium
248

When designing an ingestion strategy for a Delta Lake architecture, which TWO advantages does Delta Lake provide over traditional Parquet tables for incoming data?

Medium
249

Which security feature in Databricks allows administrators to mask sensitive data, such as email addresses or social security numbers, in query results based on user-defined functions?

Easy
250

A data engineer needs to run a Databricks notebook that processes data stored in an external ADLS Gen2 location. The notebook must access the data securely without embedding credentials in the notebook code. The engineer has already configured a Unity Catalog external location with a storage credential. Which method should the engineer use to read the data?

Easy
251

What is the primary function of an 'Access Connector' for Azure Databricks when using Unity Catalog?

Easy
252

Refer to the exhibit. A data engineer is attempting to ingest JSON data into a Delta table. Based on the error log, what is the most appropriate transformation step to implement before loading this data into the production table?

Medium
253

A junior data engineer writes a PySpark transformation that reads a Parquet dataset, filters out inactive users, and appends the resulting DataFrame to an existing Delta Lake table. However, the data engineer notices duplicate records appearing in the target table after multiple runs. Which technique should be implemented to ensure idempotency?

Easy
254

A team is using Databricks Asset Bundles to manage a job that writes to a Unity Catalog table. The bundle is deployed to a staging workspace for testing and then to a production workspace. The team wants the job to use different catalog and schema names in each environment without duplicating the entire bundle. Which approach should they use?

Hard
255

Which TWO of the following capabilities are native to the Databricks Unity Catalog?

Medium
256

Which TWO of the following are benefits of using Unity Catalog over the legacy Hive Metastore?

Medium
257

A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?

Medium
258

A data engineer is configuring Unity Catalog to allow a service principal to run a Databricks job that writes to a Delta table in an external location. The service principal has been granted USAGE on the catalog and schema, and MODIFY on the table. However, the job fails with an error indicating that the service principal cannot access the external location. What is the most likely missing privilege?

Medium
259

A data engineer is designing a Delta Lake table that will be used for both batch and streaming reads. The table must support upserts from a streaming source and maintain ACID transactions. The engineer wants to ensure that the table can be efficiently queried by downstream consumers using SQL while minimizing storage costs. Which TWO actions should the engineer take to meet these requirements? (Choose two.)

Hard
260

A data engineer needs to automate the ingestion and transformation of data within Databricks with minimal manual intervention. Which feature is most appropriate for orchestrating these data workflows?

Medium
261

A data engineer has a Delta table named silver_events with columns event_id (string), event_ts (timestamp), and payload (string). The table is partitioned by event_date (derived from event_ts). The engineer needs to update the payload column for all events that occurred on '2024-06-01' based on a mapping table named updates (event_id, new_payload). Which PySpark operation should be used to perform this update efficiently while preserving Delta Lake ACID guarantees?

Medium
262

When a Data Engineer uses a 'Repair and Rerun' functionality on a failed Databricks Job, what happens?

Medium
263

A team wants to ensure that code changes to their Databricks notebooks are reviewed before being deployed to production. They use a Git repository and Databricks Repos. Which practice should they implement in their CI/CD process?

Easy
264

Which identity management approach is mandatory for the implementation of Unity Catalog within a Databricks environment?

Medium
265

A CI/CD pipeline runs unit tests against transformation logic before deploying to production. The tests must run on a Databricks cluster and produce a pass/fail result that fails the pipeline when assertions do not hold. Which implementation best fits this requirement?

Hard
266

A data engineer needs to perform an upsert on a target Delta table using a source DataFrame. Which operation provides the most robust mechanism to handle duplicates and updates in a single pass?

Medium
267

A data engineer needs to create a Delta table in Unity Catalog that will be used by multiple teams. The engineer wants to ensure that the table is governed by Unity Catalog and that access can be controlled using SQL GRANT statements. What is the correct way to create the table?

Easy
268

Refer to the exhibit. A data engineer is deploying an automated job. Based on the provided configuration, what is the primary benefit of using the 'autoscale' attribute in this cluster definition?

Medium
269

When a data engineer needs to automate a recurring ETL job, which Databricks tool is most appropriate for orchestrating tasks and handling dependencies?

Medium
270

A data engineer needs to run a Databricks notebook on a schedule every day at 8:00 AM UTC. The notebook performs data transformations and writes results to a Delta table. The engineer wants to ensure the job runs reliably and can be monitored. Which Databricks feature should the engineer use to schedule and monitor the notebook?

Easy
271

A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?

Medium
272

A data engineer needs to ensure that only members of the 'finance' group can view a specific column containing credit card numbers in a Unity Catalog table. Which feature should be used?

Easy
273

A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?

Hard
274

Which Databricks compute resource is specifically optimized for running BI dashboards and SQL queries?

Easy
275

A data engineer maintains a Lakeflow Job with a scheduled trigger set to run every day at 08:00. The job's source table is refreshed by an upstream process that sometimes finishes later than expected, causing the job to process stale data. The engineer wants the job to start only after the upstream refresh completes, regardless of the clock time, while still preserving the existing 08:00 schedule as a fallback. Which trigger configuration should the engineer implement?

Medium
276

A data engineer needs to join a small 'lookup' table with a large 'transactions' table in a Spark job. Which transformation strategy will provide the best performance in a cluster environment?

Medium

Frequently asked questions

What does the scenario questions domain cover on the Databricks-DE-Assoc exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 276 scenario questions questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.