Courseiva

Databricks Certified Data Engineer Professional (Databricks-DE-Pro) — Questions 226–267

267 questions total · 4pages · All types, answers revealed

Page 3

Page 4 of 4

226
MCQmedium

A production Databricks workflow involves a task that runs a notebook. The notebook takes 15 minutes to finish, but the workflow is set to timeout after 10 minutes. What happens?

A.The workflow continues to run, but logs a warning about the duration.
B.The task is terminated and marked as 'Failed' by the scheduler.
C.The workflow pauses and waits for the engineer to manually approve continuation.
D.The task continues to run, but is moved to a background queue.
AnswerB

The scheduler enforces the timeout by killing the task process. Once terminated, the job workflow marks the task as 'Failed' because it did not complete successfully within the defined limits. This ensures that the system does not waste time and money on jobs that are behaving abnormally.

Why this answer

When a task exceeds its configured 'timeout' value, the Databricks scheduler forcibly terminates the task execution. This is a deliberate safety measure to prevent runaway processes from consuming cluster resources indefinitely. In production, this highlights the necessity of monitoring execution times and setting appropriate timeouts that account for normal data volume fluctuations while catching truly stuck jobs that could impact cost and resource availability.

Exam trap

Candidates often assume that a workflow timeout will gracefully cancel the notebook and mark it as 'Canceled' or 'Timed Out', overlooking that Databricks specifically categorizes exceeded task timeouts as 'Failed'.

227
MCQmedium

An organization wants to monitor and limit the spend of their Databricks SQL warehouses. Which feature is most appropriate for setting alerts when costs exceed a certain threshold?

A.Enable auto-termination on every individual query executed.
B.Use the Databricks Billing System to set a hard limit that automatically shuts down the account.
C.Implement Budget Policies in the Databricks account console to set alerts on usage.
D.Manually check the Spark UI after every query to calculate the cost.
AnswerC

Budget policies in the Databricks account console allow administrators to define budgets and receive proactive alerts. This is the official and most effective way to track consumption and manage costs across different workspaces, providing the visibility needed to prevent excessive spending before it impacts the monthly cloud bill.

Why this answer

Databricks provides budget policies and administrative controls to manage costs. Setting up budget alerts via the Databricks account console allows administrators to define spending limits and receive notifications when consumption approaches these levels. This proactive monitoring is essential for governance, ensuring that cost overruns are identified early and addressed by the relevant teams, preventing unexpected bills and promoting responsible use of compute resources across the organization.

Exam trap

Candidates often confuse cluster auto-termination policies with Unity Catalog permissions, forgetting that account-level budget policies are specifically designed for spend monitoring and alerts.

228
MCQhard

Which THREE of the following are benefits of using Delta Lake over standard Parquet files for your data lake storage?

A.ACID transactions for concurrent reads and writes.
B.Automatic file compaction without impacting reader performance.
C.Built-in support for time travel using transaction logs.
D.Faster cold-start reads by caching data on the driver.
E.Allows the use of SQL as the only way to interact with data.
AnswerA, B, C

ACID transactions ensure that concurrent operations do not result in partial data writes or inconsistent states. This is fundamental for multi-user, multi-process environments, allowing reliable reads even while writes are in progress, which is impossible with raw Parquet files that lack a unified transaction log layer for coordination.

Why this answer

Delta Lake adds a transactional layer (the log) on top of Parquet files, enabling ACID transactions, time travel, and schema enforcement. These features solve the most common challenges with raw data lakes, such as partial writes and data corruption. As a Databricks professional, you must prioritize Delta Lake to build reliable, scalable architectures that support robust data governance and high-performance analytical queries without the risks of file-based data inconsistency.

Exam trap

Candidates often confuse Delta Lake features with native Parquet capabilities, forgetting that raw Parquet lacks ACID transactions, built-in time travel, and native automatic file compaction features.

229
Multi-Selecthard

An organization is migrating to Unity Catalog and needs to secure sensitive data. Which TWO of the following statements regarding Unity Catalog security best practices are correct?

Select 2 answers
A.Assign table ownership to individual users for better tracking.
B.Use groups instead of individual users for access grants.
C.Ensure that the metastore admin has access to all data.
D.Use Service Principals for automated CI/CD job execution.
E.Public access should be granted to the root catalog.
AnswersB, D

Granting permissions to groups rather than individual users simplifies access management and reduces the risk of human error. When a new user joins a team, they automatically inherit the correct permissions by being added to the relevant group, ensuring consistent security posture across the entire data platform.

Why this answer

Effective security in Unity Catalog requires a deep understanding of object ownership and the principle of least privilege. By ensuring that objects are owned by a group rather than an individual, organizations prevent access gaps when staff turnover occurs. Additionally, using service principals for automated pipelines ensures that data access is tied to the workload rather than a user, maintaining consistent security postures across environments.

Exam trap

Test-takers frequently assume individual user accounts are acceptable for production CI/CD pipelines or object ownership, overlooking the maintenance nightmare when employees leave the organization.

230
Multi-Selecthard

Which TWO of the following are true regarding Unity Catalog's ability to govern external locations?

Select 2 answers
A.External locations require a storage credential to function.
B.Users can directly mount external locations as DBFS paths.
C.Access to external locations can be granted to users using GRANT statements.
D.External locations are automatically created for every S3 bucket in the account.
E.External locations are only supported for Delta-formatted data.
AnswersA, C

A storage credential acts as a bridge between Unity Catalog and the cloud storage provider. Without it, Unity Catalog would have no way to authenticate and access the files in the storage account on behalf of the user, making it impossible to manage external tables securely within the metastore.

Why this answer

Unity Catalog can manage access to external cloud storage locations by creating 'External Locations'. This allows administrators to grant specific permissions to users to read from or write to these storage buckets without providing them with direct cloud provider credentials. This abstraction is vital for security, as it centralizes control and auditing of data access within the platform while keeping the underlying storage infrastructure shielded from the users.

Exam trap

Candidates often assume that Unity Catalog grants access directly to the storage bucket using IAM roles. They overlook the mandatory intermediate 'storage credential' object required for this process.

231
MCQhard

A data engineer is tasked with ensuring that sensitive information in a 'customer' table is masked for all users except the 'Data_Science' group. What is the correct Unity Catalog feature to implement?

A.Create a view that performs a CASE statement to mask the column.
B.Apply a masking policy using a SQL function to the table column.
C.Use the 'DROP COLUMN' command to remove sensitive columns.
D.Encrypt the column using an external library before writing to Delta.
AnswerB

Unity Catalog supports masking policies using SQL functions. By defining a function that checks for group membership and returns either the original or masked value, you apply a central policy. This is the official and most efficient method to handle dynamic masking requirements across the entire organization.

Why this answer

Dynamic data masking in Unity Catalog allows for the creation of masking functions that return obfuscated values based on the current user's role. By assigning these functions to columns, administrators ensure that sensitive data is protected while remaining available for authorized users. This approach is highly effective for maintaining data usability without compromising the security of PII.

Exam trap

Candidates mistakenly choose physical data duplication or static table views with restricted access rather than dynamic column masking.

232
MCQmedium

A data engineer wants to monitor the data quality of a Delta table over time. Which tool is most appropriate for this task?

A.Use Databricks SQL Alerts to count total rows in the table every hour.
B.Implement DLT Expectations within your pipeline code to validate data constraints.
C.Write a Spark job that manually scans the table for NULL values and sends an email.
D.Configure a cluster-level alert to trigger when the table metadata is updated.
AnswerB

DLT Expectations are built-in features that allow you to specify data quality rules, such as null checks or value ranges, directly in your pipeline. They provide native metrics and alerting capabilities, making it easy to monitor the health of data as it passes through the system without manual oversight.

Why this answer

Delta Live Tables (DLT) Expectations are the most appropriate and integrated way to enforce and monitor data quality. By defining constraints directly within the DLT pipeline code, data engineers can automatically validate data as it arrives, generate failure logs, and even quarantine invalid records. This observability is critical for maintaining high-quality datasets in a lakehouse architecture without requiring complex, separate validation workflows for every table.

Exam trap

Test-takers frequently choose generic Spark dataframe validations or third-party tools instead of native Delta Live Tables expectations built specifically for streaming pipelines.

233
MCQmedium

An enterprise data team runs a large nightly batch job using a standard all-purpose cluster. The job frequently fails due to cloud provider spot instance pre-emptions and takes over four hours to complete. How should the engineer refactor this architecture for maximum cost efficiency and reliability?

A.Provision a larger all-purpose cluster with double the worker nodes to brute-force execution speed.
B.Convert the workload to use a Databricks Job cluster configured with spot instances and automatic fallback to on-demand.
C.Upgrade the cloud provider virtual machine family to the latest generation without changing cluster types.
D.Increase the Apache Spark executor memory fraction and decrease shuffle partition counts.
AnswerB

Job clusters are cheaper than all-purpose clusters and terminate after the run. Configuring spot instances with automatic fallback to on-demand preserves cost savings while surviving pre-emptions, directly addressing the reliability failure and the four-hour runtime.

Why this answer

Migrating the workload from an all-purpose interactive cluster to a Databricks Job cluster running on spot instances with an automatic fallback mechanism ensures cost-effective batch execution. Job clusters consume lower DBU rates than all-purpose clusters, and spot instances drastically reduce infrastructure costs while fallback guarantees completion despite cloud provider interruptions.

Exam trap

Candidates often select 'All-Purpose Clusters' for production jobs because they are easier to manage, failing to recognize that Job clusters are cheaper and more reliable for automated tasks.

234
MCQmedium

A Data Engineer is working on a Delta Live Tables (DLT) pipeline that ingests JSON files from cloud storage. The pipeline must drop rows where the 'email' column is null and also flag rows where 'age' is negative as invalid, but still process them. Which combination of DLT expectations should be used?

A.Use @dlt.expect_all({"valid_email": "email IS NOT NULL", "non_negative_age": "age >= 0"})
B.Use @dlt.expect_or_fail("valid_email", "email IS NOT NULL") and @dlt.expect_or_drop("non_negative_age", "age >= 0")
C.Use @dlt.expect_all_or_drop({"valid_email": "email IS NOT NULL", "non_negative_age": "age >= 0"})
D.Use @dlt.expect_or_drop("valid_email", "email IS NOT NULL") and @dlt.expect("non_negative_age", "age >= 0")
AnswerD

Using expect_or_drop for email nulls removes those rows, while expect for negative age records and retains them, flagging as invalid. This matches the requirement to drop rows with null email and flag negative ages without dropping them.

Why this answer

The correct approach uses expect_or_drop for the email null rule to remove those rows, and expect for the age rule to record violations without dropping. This satisfies both the drop and flag requirements simultaneously. Other expectation types either fail the pipeline, only log metrics, or drop rows that should be retained.

Exam trap

The trap here is confusing expect_or_drop with expect; the former drops rows, while the latter only records metrics.

235
Multi-Selecthard

When configuring a Service Principal to access a Unity Catalog-enabled workspace, which THREE steps are required to ensure secure and functional access?

Select 3 answers
A.Create the Service Principal in the cloud provider's IAM.
B.Assign the Service Principal the 'Owner' role on the entire workspace.
C.Grant the Service Principal access to the required Unity Catalog objects.
D.Add the Service Principal to the workspace using the SCIM API or UI.
E.Enable multi-factor authentication (MFA) for the Service Principal.
AnswersA, C, D

The Service Principal must originate in the cloud provider's IAM system, such as AWS IAM or Azure AD. This provides the identity that Databricks will use to authenticate requests, ensuring that the identity is managed and secured according to the organization's enterprise-wide security standards and policies.

Why this answer

Service Principals are the recommended way to automate CI/CD and production jobs. Configuring them correctly involves provisioning the identity in the cloud provider, syncing it with the Databricks account, and granting specific permissions on the data objects. This lifecycle management is critical for operational security and ensuring that automated processes have the least privilege required to function correctly in a production environment.

Exam trap

Test-takers frequently forget that service principals must be explicitly added to the workspace via SCIM alongside cloud IAM configuration.

236
Multi-Selectmedium

Which TWO of the following are valid ways to trigger a job in Databricks?

Select 2 answers
A.Using the 'Jobs' UI to trigger a run.
B.Executing the 'run_job' command inside a notebook cell.
C.Calling the Databricks REST API (Jobs API).
D.Running the 'dbutils.jobs.run()' method.
E.Creating a file named 'trigger.job' in DBFS.
AnswersA, C

The UI provides an easy-to-use interface for manually running a job, viewing historical execution results, and configuring schedules. This is ideal for development, testing, and troubleshooting, as it allows engineers to quickly run and monitor jobs without writing code or interacting with APIs.

Why this answer

Databricks provides multiple interfaces for job orchestration. The 'Jobs' UI in the workspace allows for interactive configuration and scheduling. Alternatively, the Databricks REST API provides a programmatic way to trigger jobs, which is crucial for CI/CD pipelines and integrating Databricks with external orchestrators like Airflow or Azure Data Factory.

Using these methods ensures that pipelines can be automated, monitored, and integrated into broader enterprise data workflows.

Exam trap

Candidates often select UI-only options while ignoring the REST API. In a professional data engineering context, automation via API is just as valid and common as manual UI triggering.

237
MCQeasy

A team has a large Delta table that is rarely updated. What is the most cost-effective way to store this data while maintaining the ability to query it with Databricks SQL?

A.Load all data into a high-performance in-memory database.
B.Store the data in Delta format on cloud object storage.
C.Replicate the data into multiple cloud regions for higher availability.
D.Convert the table to a legacy Hive format for better compatibility.
AnswerB

Object storage is highly durable and inexpensive. Delta Lake allows you to query this data with high performance using Databricks SQL while keeping the storage costs at the lowest possible tier. This is the standard, most cost-effective architecture for large, rarely updated datasets in a modern data lakehouse.

Why this answer

Storing data in Delta format on cloud object storage (like S3 or ADLS) provides the best cost-to-performance ratio. Databricks SQL can query this data directly without needing to load it into a proprietary data warehouse. By using object storage, you only pay for the storage used, and because the data is rarely updated, the overhead of maintenance is minimal, making it the ideal cost-optimized storage pattern.

Exam trap

Candidates mistakenly choose proprietary data warehouse storage tiers or caching mechanisms for cold data, driving up unnecessary infrastructure costs.

238
MCQeasy

What is the primary benefit of using a Job Cluster instead of an All-Purpose Cluster for production workloads?

A.Job clusters support more advanced libraries than All-Purpose clusters.
B.Job clusters are more cost-effective and provide better workload isolation.
C.Job clusters can be manually resized while the job is currently running.
D.Job clusters are required to access data stored in the Unity Catalog.
AnswerB

Job clusters provide significant cost savings because they are transient and terminate automatically upon task completion. They also offer better workload isolation, as they are dedicated to a specific job, preventing resource contention or performance fluctuations caused by other users running interactive queries on the same cluster.

Why this answer

Job clusters are specifically designed for automated production workloads. They are cheaper because they are ephemeral, terminating as soon as the job finishes, and they provide better isolation, ensuring that production jobs do not interfere with other development tasks. This is a critical best practice for cost management and system reliability, ensuring that production pipelines run in a clean, predictable environment that scales according to actual need.

Exam trap

Candidates often choose All-Purpose clusters for production tasks because they stay alive, failing to recognize that Job clusters offer better workload isolation and significantly lower costs.

239
MCQmedium

A Data Engineer wants to monitor cluster health proactively. Which metric is most effective for identifying that a cluster needs to be scaled up to handle increasing workload demands?

A.Query result set size
B.Cluster CPU utilization
C.Delta table file count
D.User login frequency
AnswerB

High CPU utilization is a direct indicator of compute saturation. When worker nodes sustain high CPU loads, processing throughput drops, leading to job latency. Monitoring this metric allows for the implementation of auto-scaling policies to add nodes, effectively balancing the workload across a larger compute resource pool.

Why this answer

Monitoring 'cluster_memory_utilization' or 'node_cpu_load' is essential for identifying saturation. When these metrics consistently approach high thresholds, it signals that the current compute power is insufficient for the data volume. Proactive scaling based on these metrics prevents job failures and performance degradation, ensuring that data pipelines meet their processing deadlines and maintain efficiency without wasting costs on over-provisioned infrastructure.

Exam trap

Candidates often select 'cluster logs' or 'job duration' as the primary metric. While useful, CPU and memory utilization are the direct indicators of resource saturation requiring vertical or horizontal scaling.

240
MCQmedium

A data engineer is configuring an Auto Loader stream to ingest JSON files from an S3 bucket into a Bronze Delta table. The source bucket contains both .json and .json.gz files, and the engineer wants to ensure that only .json files are processed. Which parameter should be set to achieve this?

A.cloudFiles.schemaLocation
B.cloudFiles.includeExistingFiles
C.cloudFiles.format
D.cloudFiles.pathGlobFilter
AnswerD

The cloudFiles.pathGlobFilter parameter allows you to specify a glob pattern to filter files based on their path or extension. By setting it to '*.json', Auto Loader will only process files ending with .json, excluding .json.gz files. This is the correct way to selectively ingest files by extension in Auto Loader, ensuring only the desired files are read.

Why this answer

Auto Loader provides the cloudFiles.pathGlobFilter option to filter files using glob patterns. Setting it to '*.json' ensures that only files with the .json extension are processed, while other files like .json.gz are ignored. This is the intended mechanism for file selection based on path patterns in Auto Loader, making it the correct choice for this scenario.

Exam trap

The trap here is confusing the parameter that specifies the data format with the one that filters files by extension; cloudFiles.format defines parsing, not selection.

241
Multi-Selecthard

A data engineer is optimizing a Spark job that reads from a large Delta table and performs a join with a smaller dimension table. The job is running slowly, and the engineer suspects data skew and shuffle overhead are the main issues. Which two techniques should the engineer apply to improve performance and reduce cost? (Choose two.)

Select 2 answers
A.Increase the number of shuffle partitions to 2000.
B.Cache the larger table in memory before the join.
C.Repartition the larger table on the join key before the join.
D.Enable adaptive query execution (AQE) and set spark.sql.adaptive.skewJoin.enabled to true.
E.Use broadcast join for the smaller dimension table.
AnswersD, E

Adaptive Query Execution (AQE) dynamically optimizes query plans at runtime, including handling skew by splitting skewed partitions. Enabling skew join optimization allows Spark to automatically detect and mitigate skew during joins, improving performance without manual intervention. This reduces shuffle overhead and prevents straggler tasks, leading to faster completion and lower cost.

Why this answer

Broadcast join eliminates shuffle for the smaller table, and enabling AQE with skew join optimization dynamically handles skew in the larger table. Together, they reduce shuffle overhead and mitigate stragglers, improving performance and lowering cost. Other options either introduce unnecessary shuffles or do not directly address skew.

Exam trap

The trap here is thinking that repartitioning or increasing shuffle partitions will solve skew, when in fact they can add overhead without addressing the root cause.

242
MCQmedium

A financial institution is building a Gold layer table that must support point-in-time queries to reconstruct account balances as of any past date. The source data includes transactions with effective dates and an audit log of changes. Which modeling technique is most appropriate?

A.Use Delta Lake time travel to query previous versions of the table.
B.Use a Type 1 slowly changing dimension (SCD) to overwrite old values.
C.Create a Type 3 slowly changing dimension (SCD) with previous value columns.
D.Implement a Type 2 slowly changing dimension (SCD) with effective start and end dates.
AnswerD

Type 2 SCD preserves history by creating new rows for changes, with effective start and end dates. This allows point-in-time queries by filtering on the desired date. In Databricks, this can be implemented using Delta Lake's merge operations and time travel. It is the standard technique for temporal analysis in data warehousing.

Why this answer

A Type 2 SCD retains full history by adding new rows with effective date ranges, enabling accurate point-in-time queries. This is essential for financial data where past states must be reconstructable. Delta Lake time travel is limited by retention and not designed for continuous history.

Type 1 and Type 3 SCDs lack the necessary historical depth.

Exam trap

The trap here is confusing Delta Lake time travel with a full history tracking mechanism, but time travel only retains versions for a limited period and does not model changes explicitly.

243
MCQmedium

A data engineering team stores customer transaction data in a Unity Catalog managed table named prod.finance.transactions. The security team requires that any query referencing this table, whether through a view or directly, is recorded with the identity of the user who ran it, and that the audit logs are retained for 365 days. The workspace uses Unity Catalog and has audit logs delivered to a cloud storage location. Which configuration should the data engineer verify or set to meet the requirement that all access to the table is captured with the user identity?

A.Configure a cluster policy that enforces the use of a specific Spark log4j appender to send query logs to the audit storage.
B.Enable table access control on the cluster and set the cluster's Spark configuration to log all queries.
C.Ensure that Unity Catalog audit logging is enabled for the account and that the audit log delivery is configured to the required cloud storage with appropriate retention.
D.Create a view over the table and grant SELECT on the view only, so that all access goes through the view and is logged.
AnswerC

Unity Catalog automatically records access events, including the user identity, for all queries against Unity Catalog objects. These events are written to the account-level audit log. To meet the 365-day retention requirement, the audit log must be delivered to a cloud storage location configured with the necessary lifecycle policy. No additional cluster-level setting is needed to capture the user identity.

Why this answer

Unity Catalog audit logs are generated at the account level and capture the identity of the user performing the action on Unity Catalog objects. To satisfy a retention requirement, the logs must be delivered to durable cloud storage with a retention policy. Cluster-level or view-level controls do not replace this native audit logging, and they do not provide the same structured, tamper-resistant record.

Exam trap

The trap here is assuming that cluster-level query logging or view-based access is equivalent to Unity Catalog audit logging, which already records user identity for all Unity Catalog access.

244
Multi-Selectmedium

A data engineer is setting up a Delta Share to provide an external partner with access to a subset of data. The partner will use a non-Databricks client that supports the Delta Sharing protocol. Which two actions are required to enable the partner to access the shared data? (Choose two.)

Select 2 answers
A.Create a recipient in Unity Catalog and provide the partner with the activation link or credential file.
B.Configure a Databricks SQL warehouse for the recipient to query the shared data.
C.Add the table to the share and grant the recipient access to the share.
D.Enable the Delta Sharing server on the Databricks workspace.
E.Grant the recipient SELECT privileges on the shared table.
AnswersA, C

Creating a recipient in Unity Catalog is essential to establish the sharing relationship. The recipient is issued a credential file or activation link that the partner uses to authenticate to the Delta Sharing server. Without this, the partner cannot access the share. This step is a core part of the Delta Sharing setup process, ensuring secure and controlled access.

Why this answer

To enable an external partner to access a Delta Share, two key steps are required: creating a recipient and providing them with the credential file or activation link, and adding the table to a share while granting the recipient access to that share. These actions establish the secure sharing relationship and define the data scope. Other options are either not applicable to non-Databricks clients or not part of the Delta Sharing setup process.

Exam trap

The trap here is thinking that standard Unity Catalog SQL grants like SELECT are used for Delta Sharing recipients, when access is actually controlled at the share level.

245
MCQmedium

A Spark job is failing with an OutOfMemoryError (OOM) during a group-by operation on a skewed key. What is the most effective way to resolve this?

A.Increase the executor memory size for all workers in the cluster.
B.Use the 'salting' technique by adding a random prefix to the skewed key.
C.Remove the group-by clause and process the data using a UDF.
D.Reduce the number of executors to force the job to run sequentially.
AnswerB

Salting distributes the skewed data across multiple tasks by adding a random prefix to the grouping key. This ensures that the heavy skewed key is spread out, preventing any single task from hitting memory limits. It is a robust, standard solution to handle data skew in Spark.

Why this answer

Data skew occurs when one key contains a disproportionate amount of data, causing one task to process significantly more than others. Salting (adding a random prefix to the key) breaks this concentration, distributing the data evenly across the cluster. This prevents individual tasks from crashing due to memory limits, allowing the job to complete successfully and efficiently by balancing the workload across all available workers.

Exam trap

Candidates often suggest increasing cluster size (vertical scaling) or memory settings. These are inefficient 'band-aid' fixes that do not address the root cause of uneven data distribution.

246
MCQmedium

You are writing a PySpark script to join two large tables. You want to ensure the join operation is optimized for performance by broadcasting the smaller table. Which configuration property should you adjust, or code construct should you use, to force this behavior?

A.spark.conf.set('spark.sql.shuffle.partitions', '1')
B.spark.sql('SET spark.sql.autoBroadcastJoinThreshold = -1')
C.df.join(broadcast(small_df), 'id')
D.df.repartition(100)
AnswerC

The broadcast function provides a hint to the Catalyst optimizer to broadcast the specific dataframe to all worker nodes. This is the most efficient way to ensure a broadcast join occurs regardless of the default autoBroadcastJoinThreshold configuration, effectively eliminating the need for network-heavy shuffles during the join operation.

Why this answer

Broadcasting small tables is a critical optimization technique in Databricks to prevent expensive shuffles across the cluster. By utilizing the broadcast hint, developers explicitly instruct the Spark Catalyst optimizer to send the smaller table to all worker nodes. This minimizes data movement and significantly reduces latency during join operations.

Understanding this mechanism is essential for building scalable ETL pipelines and ensuring efficient cluster resource utilization within the Databricks unified analytics platform.

Exam trap

Candidates often try to manually set 'spark.sql.autoBroadcastJoinThreshold' to a massive value, which can cause driver OOM errors, instead of using the explicit 'broadcast()' hint on the specific dataframe.

247
MCQhard

A data engineer is deploying a Databricks Asset Bundle (DAB) that defines a job with a notebook task. The bundle validates locally, but deployment fails with 'Error: cannot find notebook at path /Workspace/Users/dev@example.com/pipeline/ingest'. The engineer confirms the notebook exists in the workspace at that exact path. Which action should the engineer take to resolve the deployment failure?

A.Add the spark.databricks.workspace.path Spark config to the job cluster
B.Ensure the notebook path in the bundle configuration is relative to the bundle root and that the sync root includes the notebook
C.Convert the notebook task to a Python wheel task
D.Change the notebook task to use a Git source instead of a workspace path
AnswerB

DABs resolve notebook paths relative to the bundle root and synchronize files to the workspace. If the path is absolute or points outside the sync root, deployment cannot find it. Making the path relative and confirming the sync root includes the notebook ensures the file is uploaded and the job references the correct workspace location.

Why this answer

Databricks Asset Bundles synchronize files from the bundle root to the workspace. Notebook paths must be relative to the bundle root and included in the sync. An absolute or incorrectly rooted path causes deployment to fail even if the file exists in the workspace, because the bundle cannot map it to a synchronized file.

Exam trap

The trap here is assuming that because the notebook exists in the workspace, the bundle can reference it directly; bundles require relative paths within the sync root.

248
MCQmedium

A Data Engineer is building a Databricks SQL pipeline that ingests clickstream events from a Delta table. The events table contains a nested column `payload` of type STRUCT with fields `page_id` (STRING), `duration` (INT), and `referrer` (STRING). The engineer needs to flatten the `payload` fields into top-level columns and drop any records where `page_id` is NULL. Which SQL expression accomplishes this transformation while preserving all other columns?

A.SELECT *, payload.* AS (page_id, duration, referrer) FROM events WHERE payload.page_id IS NOT NULL
B.SELECT *, payload.page_id AS page_id, payload.duration AS duration, payload.referrer AS referrer FROM events WHERE payload.page_id IS NOT NULL
C.SELECT *, explode(payload) AS (page_id, duration, referrer) FROM events WHERE page_id IS NOT NULL
D.SELECT *, payload[0] AS page_id, payload[1] AS duration, payload[2] AS referrer FROM events WHERE payload[0] IS NOT NULL
AnswerB

This correctly uses dot notation to extract nested fields from the STRUCT column and renames them as top-level columns. The WHERE clause filters out records where the nested page_id is NULL, preserving all other columns via the wildcard. Databricks SQL supports this syntax for nested data, making it a valid and efficient solution for flattening and cleansing clickstream events.

Why this answer

The correct solution uses dot notation to access nested STRUCT fields, renames them to top-level columns, and applies a filter on the nested field to drop NULL page_id records. This is the standard and supported way in Databricks SQL to flatten and cleanse nested data while retaining all other columns. The other options misuse functions or syntax not applicable to STRUCT types.

Exam trap

The trap here is confusing STRUCT field access with array indexing or assuming that explode works on STRUCTs, when it is only for arrays and maps.

249
MCQmedium

A Data Engineer is building a Delta Live Tables (DLT) pipeline to ingest raw JSON data. They need to ensure that records missing the required 'user_id' field are dropped while simultaneously capturing these discarded records in a separate table for auditing purposes. Which approach achieves this in DLT?

A.Apply the 'expect_or_drop' constraint and use a trigger to log dropped records.
B.Use the 'expect_or_fail' constraint to halt the pipeline and flag the error.
C.Define two separate tables in the pipeline, one filtering for valid records and one for invalid records using the NOT condition.
D.Configure a DLT pipeline to use 'expect_all_drop' to automatically split the data stream.
AnswerC

By defining two tables, you leverage the declarative nature of DLT to materialize valid and invalid datasets concurrently. Using the inverse logical condition for the audit table ensures that all records are accounted for, meeting both the ingestion requirement and the audit policy without halting the pipeline's progress.

Why this answer

To handle data quality in DLT, the 'expect_violation_or_drop' constraint is not a standard clause. Instead, the 'expect_or_drop' constraint removes invalid records, but does not preserve them. The correct architectural pattern involves using a separate pipeline or query that filters for the inverse condition (where 'user_id' is null) and writes those records to an 'expect_all_fail' target table, maintaining strict lineage and auditing for schema non-compliance.

Exam trap

Candidates often search for a single DLT command that drops and saves records simultaneously. They fail to realize that DLT requires two distinct logic paths to separate valid and invalid data.

250
Multi-Selectmedium

A data engineer is designing a Unity Catalog governance model for a new data lakehouse. They need to ensure that data access is auditable and that sensitive data is protected. Which two actions should the engineer take to meet these requirements? (Choose two.)

Select 2 answers
A.Enable audit logging for the metastore to capture all access and permission changes.
B.Use dynamic views to mask PII columns for all users except administrators.
C.Apply tags to sensitive columns and use attribute-based access control (ABAC) policies to restrict access.
D.Create a separate metastore for each department to isolate data.
E.Store all data in a single catalog and grant `ALL PRIVILEGES` to the data engineering team.
AnswersA, C

Audit logging in Unity Catalog records detailed events such as data access, permission changes, and metadata operations. Enabling it provides the necessary audit trail to track who accessed what data and when, which is essential for compliance and security monitoring. This directly addresses the requirement for auditable data access.

Why this answer

Enabling audit logging captures all access and permission changes, providing the necessary audit trail. Applying tags to sensitive columns and using ABAC policies enforces fine-grained access control based on those tags, protecting sensitive data consistently. Together, these actions create a governance model that is both auditable and secure, leveraging Unity Catalog's native capabilities for centralized policy management.

Exam trap

The trap here is assuming that broad privileges or separate metastores provide security and auditability, when in fact they increase risk and complexity, while overlooking the native audit logging and tag-based ABAC features.

251
MCQeasy

A data engineer is configuring a Databricks Workflow that must run a notebook task only after a previous task that writes to a Delta table has completed successfully. The engineer wants to ensure that if the first task fails, the second task does not run. Which feature should the engineer use to define this dependency?

A.Configure the second task with a run_if condition set to ALL_SUCCESS.
B.Use a trigger to start the second task when the first task completes, using a file arrival trigger on the Delta table location.
C.Add a condition task between the two tasks that checks the Delta table for new records and only proceeds if records exist.
D.Set the depends_on property of the second task to reference the first task's task_key.
AnswerD

The depends_on property in a Workflow task definition specifies upstream tasks that must complete successfully before the dependent task starts. By referencing the first task's task_key, the second task will only run if the first succeeds, satisfying the requirement.

Why this answer

Task dependencies in Databricks Workflows are defined using the depends_on property, which lists upstream task keys. This ensures the dependent task runs only after the specified upstream tasks complete successfully. It is the fundamental way to build a directed acyclic graph (DAG) of tasks within a job.

Exam trap

The trap here is confusing run_if conditions with task dependencies; run_if controls behavior based on overall job status, not on specific upstream tasks.

252
MCQmedium

A data engineer is setting up a new Unity Catalog metastore. What is the primary purpose of the 'Metastore Admin' role?

A.To create and manage compute clusters for all users.
B.To define the root storage location and manage top-level catalogs.
C.To monitor and respond to daily system health alerts.
D.To approve all requests for new user sign-ups.
AnswerB

The Metastore Admin has the authority to configure the root storage location and create catalogs, which are the containers for all other data objects. This role is responsible for the overall hierarchy and security structure of the data estate, making it the most critical role for initial setup.

Why this answer

The Metastore Admin is the highest-privilege role in Unity Catalog, responsible for the initial configuration and top-level governance settings. This role is essential for establishing the security foundation of the platform, including catalog creation and cross-workspace access control. Proper management of this role is critical to prevent privilege escalation and ensure that only authorized personnel can define the organization's overall data governance strategy.

Exam trap

Candidates often confuse the Metastore Admin with a Workspace Admin. They assume the role manages individual workspace settings rather than the overarching metastore-level governance and root storage configurations for the entire account.

253
MCQmedium

An engineer is writing a Python function to process data in a Databricks Notebook. Which command should they use to ensure that secrets, such as API keys, are not hardcoded or exposed in the plain text of the notebook?

A.os.environ.get('API_KEY')
B.dbutils.secrets.get(scope='my_scope', key='my_key')
C.open('/secret/path').read()
D.spark.conf.get('secret_key')
AnswerB

This is the correct function to retrieve a secret value securely. The value returned by this function is automatically redacted if printed in the notebook logs, providing a layer of protection against accidental exposure. It integrates directly with Databricks Secret Scopes, ensuring centralized management and controlled access to sensitive credentials used in code.

Why this answer

The dbutils.secrets.get() utility is the standard, secure way to retrieve sensitive information stored in Databricks Secret Scopes. By referencing the scope and key name, the secret value is fetched at runtime and remains masked in the notebook output, preventing accidental exposure of credentials. This is a mandatory practice in any secure data engineering environment to comply with security policies and prevent unauthorized access to downstream data sources.

Exam trap

Candidates often suggest using environment variables or hardcoded strings, failing to realize these are easily exposed in notebook logs or version control, violating security best practices.

254
MCQmedium

A Data Engineer needs to ensure that PII data in a Delta table is accessible only to members of the 'hr_admin' group, while allowing all other users to view the non-PII columns. Which Unity Catalog feature is the most efficient way to implement this requirement?

A.Create separate physical tables for HR and general users.
B.Use standard SQL views for every user to filter columns.
C.Apply a column mask using a SQL function in Unity Catalog.
D.Assign the 'SELECT' permission on individual columns in the UI.
AnswerC

Column masking in Unity Catalog allows administrators to define functions that dynamically redact or obscure data based on the user's role. This provides a unified, policy-driven approach to data security that is applied at query time, ensuring compliance without the complexity of managing numerous views or physical tables.

Why this answer

Unity Catalog's row-level security and column-level masking allow for fine-grained access control directly at the table level. By defining a masking function or a column filter, the Data Engineer ensures that the security policy is enforced consistently across all SQL warehouses and Databricks Runtime versions. This approach centralizes governance, reduces administrative overhead compared to view-based security, and ensures that data privacy compliance is maintained without duplicating data or creating multiple table versions.

Exam trap

Candidates frequently suggest creating duplicate filtered views or tables for different user groups, ignoring Unity Catalog's modern column masking capabilities which provide centralized efficiency.

255
MCQeasy

A Data Engineer is using Lakehouse Federation to query a Snowflake database from Databricks. The engineer has created a connection and a foreign catalog. Which statement correctly describes how data is accessed when a user queries a table in the foreign catalog?

A.The data is cached in the Databricks workspace after the first query.
B.The data is copied into Delta Lake before the query runs.
C.The query is executed on Databricks, and the data is streamed from Snowflake.
D.The query is executed on Snowflake, and only the results are returned to Databricks.
AnswerD

Lakehouse Federation pushes down queries to the external database when possible. For Snowflake, the query is executed on Snowflake, and only the result set is returned to Databricks. This minimizes data movement and leverages Snowflake's compute. The foreign catalog provides a unified interface, but the execution happens remotely.

Why this answer

Lakehouse Federation enables querying external databases without moving data. When a user queries a foreign catalog table, Databricks pushes the query down to the external database, such as Snowflake. The external database executes the query and returns only the results.

This approach minimizes data transfer and leverages the external system's compute. The foreign catalog provides a seamless experience, but the data remains in the source system.

Exam trap

The trap here is assuming that federation copies or caches data in Databricks, when it actually pushes down queries to the external system.

256
MCQhard

Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?

A.SELECT COUNT(*) FROM my_table WHERE updated_at < now() - interval 1 hour
B.SELECT CASE WHEN max(updated_at) < now() - interval 1 hour THEN 1 ELSE 0 END
C.SELECT updated_at FROM my_table ORDER BY updated_at DESC LIMIT 1
D.SELECT datediff(now(), max(updated_at)) FROM my_table
AnswerB

This query returns 1 if the latest data is older than one hour, allowing an alert threshold of '> 0' to trigger. This is a standard pattern for binary state alerting where the result set needs to indicate a violation condition clearly.

Why this answer

To detect stale data, the alert must compare the latest record timestamp against the current time. If the latest 'updated_at' value is less than the threshold (one hour ago), it signifies that no data has arrived in the last hour. Monitoring data freshness is critical for time-sensitive analytics, ensuring that business stakeholders are working with current data and allowing teams to troubleshoot ingestion pipeline failures proactively.

Exam trap

Candidates often struggle with SQL syntax for time intervals, accidentally choosing options that use incorrect date functions or fail to account for the 'CASE WHEN' logic required for binary alert triggers.

257
MCQmedium

Refer to the exhibit. You are appending data to an existing Delta table. What is the most likely cause of this error, and how should you resolve it?

A.The table is locked by another process.
B.The incoming data's 'price' column has a higher precision than the table schema.
C.The table schema is corrupted and needs to be repaired.
D.The user does not have write access to the table.
AnswerB

The error explicitly states an incompatibility between the two decimal types. Appending data requires the incoming schema to be compatible with the target. If the incoming 'price' requires more precision than the existing column allows, the append operation is rejected to preserve the integrity of the existing data stored.

Why this answer

This error occurs because of a data type mismatch between the incoming DataFrame and the existing table schema, specifically a change in precision or scale. Databricks enforces schema safety to prevent data corruption. Resolving this requires either casting the incoming data to match the target schema or using the 'mergeSchema' option if the goal is to allow evolution, provided the change is safe and intended for the application.

Exam trap

Candidates often assume the error is due to a missing column. They overlook that Delta Lake enforces strict schema types and precision, and incoming data must match the defined target schema.

258
MCQeasy

A data engineer needs to grant a group of users the ability to run a specific Databricks job but not modify its configuration. The job is managed by a service principal. Which permission level should be assigned to the group on the job?

A.Can View
B.Can Run
C.Can Manage
D.Can Attach To
AnswerB

Can Run allows users to trigger the job and view its runs but not edit its configuration. This exactly matches the requirement: the group can execute the job but cannot change its settings. It is the appropriate permission level for operators who need to run jobs without administrative privileges.

Why this answer

Databricks job permissions include Can View, Can Run, Can Manage, and sometimes Is Owner. Can Run permits users to execute the job and view run results but prevents them from editing the job's configuration. This aligns with the requirement to allow running without modification.

Can View is too restrictive, while Can Manage grants excessive privileges. Can Attach To is unrelated to job permissions.

Exam trap

The trap here is confusing job permissions with cluster permissions, or assuming that Can View also allows running.

259
MCQhard

A Data Engineer is using Lakehouse Federation to query an external PostgreSQL database from Databricks. The engineer creates a foreign catalog named 'pg_catalog' using a connection that specifies the host, port, and credentials. Users report that queries against the foreign catalog fail with a permission error, even though the connection works when tested. The engineer confirms that the connection has the correct credentials and that the PostgreSQL user has SELECT privileges on the required tables. What is the most likely cause?

A.The PostgreSQL database is not in the same region as the Databricks workspace.
B.The users have not been granted USE on the foreign catalog and SELECT on the tables.
C.The connection uses a service principal instead of a user account.
D.The foreign catalog was created without specifying the 'postgresql' database type.
AnswerB

In Lakehouse Federation, after creating a foreign catalog, users must be granted USE on the catalog and SELECT on the tables within it. Without these Unity Catalog privileges, queries will fail with permission errors, even if the underlying PostgreSQL user has access. The connection credentials are used by Databricks to access the external database, but users still need Unity Catalog permissions.

Why this answer

In Lakehouse Federation, the connection stores credentials to access the external database, but end users must still have Unity Catalog privileges to query the foreign catalog. Specifically, they need USE on the foreign catalog and SELECT on the tables. Without these, queries fail with permission errors.

The connection test succeeds because it uses the stored credentials directly, bypassing user-level Unity Catalog checks. The correct answer addresses the missing Unity Catalog grants.

Exam trap

The trap here is assuming that if the connection works and the external user has SELECT, end users can query without additional Unity Catalog privileges.

260
MCQmedium

A Data Engineer is building a Lakeflow Spark Declarative Pipelines pipeline that ingests JSON sensor events. The pipeline must drop records where the `sensor_id` is NULL, ensure that `event_time` is not in the future, and continue processing without failing the update. Which combination of expectations should be used?

A.Use `@dp.expect_all({"valid_sensor": "sensor_id IS NOT NULL", "valid_time": "event_time <= current_timestamp()"})` on the dataset.
B.Use `@dp.expect_or_fail("valid_sensor", "sensor_id IS NOT NULL")` and `@dp.expect_or_drop("valid_time", "event_time <= current_timestamp()")` on the dataset.
C.Use `@dp.expect_all_or_fail({"valid_sensor": "sensor_id IS NOT NULL", "valid_time": "event_time <= current_timestamp()"})` on the dataset.
D.Use `@dp.expect_or_drop("valid_sensor", "sensor_id IS NOT NULL")` and `@dp.expect_or_drop("valid_time", "event_time <= current_timestamp()")` on the dataset.
AnswerD

These expectations drop records that violate the conditions while allowing the pipeline update to complete successfully. `expect_or_drop` is the correct decorator when you want to discard bad records and not fail the pipeline. Both conditions are expressed as SQL expressions, which are evaluated per row. The pipeline continues processing, and dropped records are tracked in event logs and metrics.

Why this answer

The pipeline must drop records that violate the conditions and continue processing. `expect_or_drop` is designed for this: it discards invalid rows and allows the update to succeed. Using `expect_or_fail` would halt the pipeline, and `expect_all` would not remove the bad records, leaving them in the target table.

Exam trap

The trap here is confusing `expect_all` (which only logs metrics) with `expect_or_drop` (which actually removes records), or assuming that `expect_or_fail` is needed to enforce data quality.

261
MCQhard

A data engineer is troubleshooting a Databricks SQL query that occasionally fails with 'Query exceeded the maximum allowed execution time' on a shared SQL warehouse. The query is a complex aggregation over a large Delta table. The engineer needs to identify the root cause and ensure the query can complete successfully. Which action should the engineer take first?

A.Increase the SQL warehouse size to provide more compute resources for the query.
B.Set the Spark configuration 'spark.sql.adaptive.enabled' to false to disable adaptive query execution.
C.Examine the query profile in the Databricks SQL query history to identify stages with high data skew or spill.
D.Change the query to use a larger cluster by switching from a SQL warehouse to an all-purpose cluster.
AnswerC

The query profile provides detailed execution metrics, including time spent per stage, data skew, and spill to disk. These insights help pinpoint why the query exceeds the time limit, such as an inefficient join or skewed data distribution. Addressing these issues can allow the query to complete within the limit.

Why this answer

The query profile in Databricks SQL provides detailed execution metrics that can reveal the root cause of long-running queries, such as data skew or spill. By analyzing the profile first, the engineer can make targeted optimizations, such as repartitioning or rewriting the query, which may resolve the timeout without unnecessarily increasing compute resources.

Exam trap

The trap here is jumping to scaling up the warehouse, which may mask the symptom but not fix the underlying inefficiency that the query profile would reveal.

262
Multi-Selectmedium

A data engineer is preparing to deploy a production Databricks Workflow that must be maintainable and auditable. The engineer wants to ensure that changes to the workflow are tracked and that failures can be diagnosed quickly. Which TWO practices should the engineer implement? (Choose two.)

Select 2 answers
A.Manually document each deployment in a shared spreadsheet
B.Use an all-purpose cluster instead of a job cluster for the workflow
C.Configure email notifications for job success and failure
D.Store the workflow definition in a Git repository and deploy using Databricks Asset Bundles
E.Enable job run logs to be delivered to a cloud storage location for long-term retention
AnswersD, E

Version-controlling workflow definitions in Git and deploying via Databricks Asset Bundles provides change tracking, code review, and reproducible deployments. This practice ensures that every modification is auditable and that the deployed workflow matches the reviewed source, which is essential for maintainability and compliance in production environments.

Why this answer

Version-controlling workflow definitions in Git and deploying with Databricks Asset Bundles ensures changes are tracked and reproducible. Delivering job logs to cloud storage retains diagnostic data for troubleshooting and auditing. Together, these practices provide the maintainability and auditability required for production workflows.

Exam trap

The trap here is choosing notification or manual documentation practices that seem helpful but do not provide the version control and durable logging needed for maintainability and auditing.

263
MCQhard

Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?

A.When the query returns a row count greater than 3600.
B.When the 'duration' column value in the query output is strictly greater than 3600.
C.When the query execution takes longer than 3600 seconds to run.
D.When the query fails and returns an error code equal to 3600.
AnswerB

The 'op' field is set to '>', and the column is 'duration' with a threshold of '3600'. Therefore, any value exceeding 3600 returned by the query result set will satisfy the condition and transition the alert to the 'Triggered' state.

Why this answer

The alert is configured to monitor the 'duration' column from the specified query result. If the returned value exceeds 3600, the alert enters the 'Triggered' state. Understanding this threshold-based mechanism is vital for maintaining pipeline performance.

If queries consistently exceed this time, the alert notifies the engineer to investigate potential resource contention, suboptimal query plans, or data volume spikes, ensuring the stability and timely execution of data processing tasks.

Exam trap

Candidates often misinterpret the alert condition logic, confusing the 'Triggered' state with the 'OK' state or failing to distinguish between the 'threshold' value and the 'column' being monitored.

264
MCQeasy

A data engineer is using Delta Sharing to share a table with an external partner. The partner needs to access the shared data using their own Databricks workspace. Which protocol does Delta Sharing use to enable this cross-platform sharing?

A.SFTP transfer of Parquet files
B.JDBC/ODBC connection to the provider's SQL warehouse
C.REST API with pre-signed URLs
D.Direct access to the provider's cloud storage with IAM roles
AnswerC

Delta Sharing uses a REST API to provide access to shared data. The recipient's client authenticates and requests data, and the server returns pre-signed URLs pointing directly to cloud storage. This allows the recipient to download the data without needing direct access to the provider's cloud storage or Databricks workspace, enabling secure cross-platform sharing.

Why this answer

Delta Sharing is an open protocol that uses a REST API to facilitate data sharing. The provider's Delta Sharing server authenticates requests and returns pre-signed URLs that allow the recipient to download data directly from cloud storage. This design enables secure, cross-platform sharing without exposing the provider's storage credentials or requiring the recipient to use Databricks.

Exam trap

The trap here is assuming that Delta Sharing uses traditional database connectivity like JDBC/ODBC or direct storage access, when it actually leverages a REST API with pre-signed URLs for secure, temporary access.

265
MCQhard

A data engineer is troubleshooting a production Databricks job that intermittently fails with 'SparkOutOfMemoryError'. The job processes large datasets with skewed partitions. The engineer wants to monitor the job to proactively detect memory pressure before failures occur. Which metric should the engineer monitor on the driver and executor nodes?

A.JVM heap memory usage
B.Disk I/O read/write throughput
C.CPU utilization percentage
D.Network throughput between executors
AnswerA

JVM heap memory usage on driver and executor nodes directly reflects memory pressure. When heap usage approaches the maximum, Spark may spill to disk or throw OutOfMemoryError. Monitoring heap usage allows the engineer to detect when memory is nearly exhausted and take action, such as increasing memory or repartitioning data, before job failure occurs.

Why this answer

JVM heap memory usage is the most direct indicator of memory pressure on Spark driver and executor nodes. By monitoring heap usage, engineers can detect when memory is approaching limits and intervene before OutOfMemoryError occurs. Other metrics like disk I/O, network throughput, and CPU utilization provide complementary information but do not directly measure memory consumption.

Exam trap

The trap here is focusing on symptoms like disk spilling or high CPU rather than the root cause—JVM heap memory usage—which directly signals memory pressure.

266
MCQhard

Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data ingestion over standard Structured Streaming pipelines?

A.DLT supports significantly higher throughput than Structured Streaming.
B.Declarative pipeline management and automated dependency handling.
C.Built-in data quality monitoring with Expectations.
D.DLT is the only way to read from cloud object storage.
E.DLT supports non-Delta storage formats for all outputs.
AnswerB, C

DLT allows you to define the pipeline in a declarative way, where the system automatically manages the creation and execution of the Directed Acyclic Graph (DAG) of dependencies. This eliminates the manual effort of coordinating complex streams, reducing operational overhead and the likelihood of human error in pipeline configuration.

Why this answer

DLT simplifies the operational complexity of data pipelines by providing automatic infrastructure management and built-in quality controls. It abstracts the configuration required for managing checkpoints, scaling, and handling schema drift, allowing engineers to focus on defining transformations. The 'Expectations' framework and declarative pipeline management provide superior observability and data quality enforcement compared to manual, imperative coding in standard Spark Structured Streaming.

Exam trap

Test-takers often confuse basic streaming features with DLT enhancements, overlooking declarative management and built-in quality expectations unique to DLT.

267
MCQmedium

A data engineer is configuring a Unity Catalog metastore to use a customer-managed key (CMK) for encryption at rest. The engineer has created the necessary Key Vault and key in Azure. Which additional configuration is required to enable CMK for the metastore?

A.Create a private endpoint for the Key Vault.
B.Assign the Key Vault Crypto Service Encryption User role to the Databricks managed identity used for the metastore.
C.Grant the Databricks access connector the Key Vault Crypto Service Encryption User role.
D.Enable soft delete on the Key Vault.
AnswerB

To use CMK for Unity Catalog metastore encryption, you must grant the Databricks managed identity (the one associated with the metastore) the Key Vault Crypto Service Encryption User role on the Key Vault. This allows Databricks to use the key for encrypting the metastore's data.

Why this answer

Unity Catalog CMK for Azure requires that the Databricks managed identity for the metastore has the Key Vault Crypto Service Encryption User role on the Key Vault. This role allows Databricks to perform wrap and unwrap operations with the key. Without this role assignment, the metastore cannot use the CMK for encryption.

Exam trap

The trap here is confusing the access connector identity with the metastore managed identity, or assuming network configurations like private endpoints are required for CMK.

Page 3

Page 4 of 4

All pages