Databricks-DE-Assoc · domain
scenario questions
Practise Databricks Certified Data Engineer Associate scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.
Focused practice
Practice scenario questions questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about scenario questions
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How the topic appears in realistic exam-style scenarios.
Which detail in the question changes the correct answer.
How to eliminate plausible but wrong options.
How to connect the question back to the wider exam objective.
Watch out for
Common scenario questions exam traps
- ▸Answering from memory before reading the full scenario.
- ▸Missing a constraint such as cost, availability, security, scope or command context.
- ▸Choosing a broad answer when the question asks for the most specific fix.
- ▸Ignoring why the wrong options are tempting.
Question index
All scenario questions questions (276)
Click any question to see the full explanation, or start a practice session above.
A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?
Medium2Which of the following is a fundamental principle of implementing effective CI/CD for Databricks workflows?
Medium3A data engineer is tasked with securing sensitive PII data in Unity Catalog. Which THREE actions are recommended to ensure robust security and compliance?
Medium4Which of the following describes the correct order of operations to configure a new external location in Unity Catalog?
Hard5A data engineer is designing a Delta Live Tables (DLT) pipeline. They need to ensure that records with missing values in the 'customer_id' column are dropped during the ingestion process. Which constraint syntax should be used?
Medium6A data engineer is preparing a notebook that must authenticate to cloud storage using a short-lived token that is automatically rotated by Databricks and is never written into the notebook source. The engineer wants the least administrative overhead while keeping secrets out of the code. Which approach should the engineer use?
Medium7Your organization requires that all data processing pipelines enforce a strict schema to prevent corrupt data from landing in the Silver layer. Which feature should be configured to ensure that only data matching the expected schema is written?
Medium8A CI/CD pipeline uses the Databricks CLI to deploy a job to production. The pipeline must ensure that the job configuration is identical across environments except for the cluster size, which differs between staging and production. Which approach best supports this requirement?
Hard9Refer to the exhibit. A security administrator applies this Unity Catalog policy. What is the impact on users in the 'analyst-group'?
Hard10Which THREE of the following are core components of the Databricks Intelligence Platform?
Medium11A data engineer notices that a production Delta Lake table is experiencing slow read performance during concurrent write operations. The table contains millions of small files. Which action should the engineer take to resolve this performance degradation?
Medium12A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)
Hard13What is the primary benefit of using Data Live Tables (DLT) for managing dependencies between tables in a pipeline?
Medium14A data engineering team uses Databricks Asset Bundles (DABs) to manage a job that must deploy to both a staging and a production workspace. The team wants to avoid hardcoding workspace-specific values such as the cluster ID and the storage path in databricks.yml. Which approach should they use?
Medium15Which feature in Unity Catalog is primarily used to track data movement and transformation history for compliance and auditing?
Easy16When considering the Databricks Intelligence Platform, what is the primary role of the 'Lakehouse' architecture?
Medium17A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?
Medium18A data engineer has a Delta table named `sales` with columns `sale_id`, `customer_id`, `amount`, and `sale_date`. They need to create a new table that contains only the `customer_id` and the total `amount` per customer for all sales in 2023. Which SQL statement correctly creates this aggregated table?
Medium19A platform team is setting up a CI/CD pipeline that deploys Databricks jobs and notebooks from a Git repository. They must ensure deployments are secure and auditable. (Choose two.)
Hard20A data engineering team is migrating a legacy data warehouse to Databricks. They want to ensure that their raw data is ingested into a 'Bronze' table in its original format. Which approach is most recommended for this ingestion layer?
Medium21Which TWO of the following scenarios are valid use cases for utilizing Delta Lake's Change Data Feed (CDF)?
Medium22A data engineer is configuring a Databricks job that must run on a schedule and send an email notification if the job fails. They want to minimize manual intervention. Which feature should they use to define the schedule and failure notification?
Medium23A data engineer needs to run a nightly transformation that reads a large Parquet dataset, writes a curated Delta table, and then immediately runs OPTIMIZE and VACUUM on that table. The engineer wants each step to be observable, retryable, and to avoid data loss if VACUUM fails. Which orchestration approach best meets these requirements?
Hard24A data engineer is building a Databricks job that processes millions of small JSON files landed in cloud storage each hour. The job currently spends most of its runtime on file listing and task scheduling overhead. The engineer wants to improve throughput without changing the downstream table schema. Which change should be made to the ingestion step?
Medium25A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save costs. Which Trigger should be used in the Structured Streaming code?
Medium26Refer to the exhibit. A data engineer is running these commands before performing heavy write operations into a Delta table. What is the primary benefit of enabling these configurations?
Hard27A data engineering team needs to restrict access to a sensitive customer table in Unity Catalog so that junior data engineers can only view non-PII columns, while senior engineers can view all columns. Which approach should be used to implement this requirement securely and efficiently?
Medium28A data engineer is using Structured Streaming to ingest data from a Kafka topic. They want to ensure that if the pipeline fails, it can resume exactly where it left off, without processing duplicate data. Which component enables this functionality?
Medium29Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?
Hard30A data engineer is configuring Unity Catalog governance for a multi-department organization. Which TWO actions require the metastore admin or catalog owner to have a workspace-independent metastore assigned? (Choose TWO)
Hard31Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?
Medium32A data engineer is migrating legacy tables to Unity Catalog. Which TWO of the following statements regarding the transition to three-level namespace (catalog.schema.table) are true?
Hard33When ingesting data from cloud storage, what is the most important reason to use a Service Principal or IAM Role instead of individual user credentials?
Medium34A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?
Hard35A data engineer is configuring a Lakeflow Job that processes sensitive customer data. The job must notify the on-call team when a run fails and must also capture the run's output for auditing. Which TWO actions should the engineer take in the Lakeflow Jobs configuration? (Choose two.)
Medium36Which security feature should a data engineer configure to ensure that audit logs from all Databricks workspaces are captured and stored in a single, centralized location?
Medium37A data engineer is using Auto Loader to ingest data from a Kafka topic into a Delta table. The engineer wants to ensure that the ingestion handles late-arriving data and provides exactly-once semantics. Which combination of features should the engineer use?
Hard38A data engineering team is building a medallion architecture in Databricks. In the Silver layer, streaming data from Kafka must be cleaned, deduplicated, and written into a Delta table. Which Structured Streaming output mode should the engineer select to ensure append-only storage of fully processed, stateful deduplicated records?
Medium39Refer to the exhibit. A data engineer is reviewing the configuration of a Delta table that is frequently queried by 'customer_id'. Given the current metadata, which action will provide the most significant improvement to query performance?
Medium40A data engineer needs to store structured data in a cloud object storage location while maintaining full ACID guarantees. Which storage format is the foundation of the Databricks Lakehouse architecture that enables this functionality?
Easy41A data engineer has a Lakeflow Job with two tasks: Task1 and Task2. Task2 must run only if Task1 succeeds. The engineer also wants Task2 to be skipped if Task1 fails, but the overall job status should be marked as failed. Which configuration should the engineer use for the dependency between Task1 and Task2?
Hard42A data engineer is configuring a Databricks cluster to run a Spark job that processes large datasets. The job requires high memory and will run for several hours. The engineer wants to minimize costs while ensuring the job completes successfully. Which cluster configuration should the engineer choose?
Medium43A data engineer needs to pass the execution date to a job task dynamically. Which feature should they use?
Medium44Refer to the exhibit. A data engineer attempts to run the GRANT statement above in a Databricks SQL query editor. What is the most likely reason for failure?
Medium45A team is setting up a CI/CD pipeline that deploys Databricks Asset Bundles to production. They want the pipeline to be secure and to fail fast before any production resources are changed. Which TWO practices should they implement? (Choose two.)
Medium46A data engineer is working with a large, partitioned table and needs to perform a complex transformation. Which THREE of the following strategies will optimize query performance for this transformation?
Hard47When promoting code from a development workspace to a production workspace, what is the primary risk of using manual notebook exports?
Medium48A data engineer needs to create a Silver Delta table that contains only distinct, non-null `customer_id` values from a Bronze table, and the result must be refreshed idempotently each night. Which statement best satisfies the requirement?
Easy49Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?
Medium50A data engineer is working on a Bronze-to-Silver transformation. Which THREE of the following practices are recommended to optimize performance and data quality during this stage?
Medium51Refer to the exhibit. A user who is a member of 'data_analysts_group' reports they cannot see the 'orders' table in the Catalog Explorer. What is the most likely cause?
Medium52What is the primary purpose of a 'feature branch' in a Git-based workflow for Databricks?
Easy53A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?
Medium54A data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?
Medium55A data engineering team is setting up a CI/CD pipeline for Databricks notebooks using Databricks Repos and a Git provider. They want to ensure that changes are tested before being merged and that production deployments are controlled. Which TWO practices should they implement? (Choose two.)
Hard56A data engineer is creating a Lakeflow Job that must run a Python script stored in DBFS. The engineer wants to ensure the script is executed with the correct dependencies and environment. Which task type should be used?
Easy57A data engineer needs to share a subset of a table with an external partner who does not have access to the internal Databricks workspace. What is the most appropriate method to achieve this?
Medium58A data engineer is working with a Delta table that contains a column 'timestamp' of type timestamp. The table is partitioned by date. The engineer needs to run a query that filters on a specific date range and also on a high-cardinality column 'user_id'. The query is performing poorly. Which optimization technique should the engineer apply to improve query performance?
Hard59A data engineer is designing a secure architecture using the Databricks Intelligence Platform. Which TWO of the following statements accurately describe the role and capabilities of Unity Catalog within this platform? (Choose TWO)
Hard60A data engineer is building a Lakeflow Job that must process a parameterized date range. The engineer wants to pass start_date and end_date values into a notebook task at runtime and have those values available as widget-like parameters inside the notebook. Which approach should the engineer use?
Medium61A data engineer is configuring audit logging for a Databricks workspace that uses Unity Catalog. The security team requires that all access to data in Unity Catalog be logged and available for analysis in a centralized location. The data engineer wants to enable the delivery of audit logs to an AWS S3 bucket. Which configuration should the data engineer use?
Medium62A data engineer is using Auto Loader to stream data from Kafka into a Delta table. The Kafka topic receives messages in Avro format, and the schema is stored in a Confluent Schema Registry. The engineer wants Auto Loader to automatically fetch the schema from the registry and evolve it as new versions are registered. Which configuration should the engineer use?
Hard63Refer to the exhibit. A job fails with the provided error message. What is the most likely cause of this failure in a Databricks Delta Lake environment?
Hard64An organization is adopting the Databricks Intelligence Platform and wants to leverage Mosaic AI for building custom machine learning models. Which feature allows data engineers to track machine learning experiments, log parameters, and manage model artifacts reliably?
Medium65A data engineer is configuring a Databricks Auto Loader stream to ingest CSV files from a cloud storage location into a Delta table. The CSV files have a header row, and the engineer wants to automatically infer the schema and store the inferred schema in a specified location for consistency across restarts. Which Auto Loader option should be used to persist the inferred schema?
Medium66A data engineer is managing a Unity Catalog table that contains sensitive financial data. The table is owned by the 'finance' group, and the data engineer needs to allow the 'auditors' group to read the table but not modify it. Additionally, the data engineer wants to ensure that the 'auditors' group can see the table's metadata (e.g., column names and types) but cannot access the underlying data files directly. Which TWO actions should the data engineer take to meet these requirements? (Choose two.)
Medium67A data engineer needs to allow a service principal to read data from a Unity Catalog table `sales.orders` and also write to a volume `sales.raw_data`. Which set of privileges should be granted to the service principal?
Medium68A data engineering team runs a nightly Lakeflow Job that ingests files from cloud storage, transforms them with a notebook, and then runs a SQL task. The team wants the SQL task to execute only after the notebook transform succeeds, but they do not want the SQL task to wait for a fixed delay. Which Lakeflow Jobs feature should they configure on the SQL task?
Medium69A data engineer needs to grant the `analyst` group the ability to read data from a Unity Catalog table `sales.fact_orders`. They also want to ensure that members of `analyst` can see the table in the catalog explorer but cannot modify it. Which privilege should be granted to the `analyst` group on the table?
Easy70A team is designing a CI/CD pipeline using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for managing Databricks projects?
Medium71A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?
Medium72A data engineer is designing a Bronze-to-Silver transformation pipeline using Delta Lake. They need to ensure that the Silver table contains only records where the 'transaction_id' is not null and the 'amount' is positive. Which technique best ensures data quality at this stage?
Medium73A job is failing with a 'Disk Space' error on the worker nodes. The code performs several large joins. Which configuration should the engineer adjust to mitigate the disk space usage?
Medium74When implementing a Medallion Architecture, what is the primary purpose of the 'Silver' layer?
Easy75A data engineer is using Delta Live Tables to build a pipeline. They need to create a table that contains the latest record for each customer based on a `last_updated` timestamp. The source is a streaming table with append-only data. Which Delta Live Tables operation should be used to achieve this?
Medium76Which THREE of the following are benefits of using Delta Live Tables (DLT) for managing your data pipelines?
Hard77A data engineer is using Auto Loader and wants to handle a situation where a column 'user_id' is sometimes an integer and sometimes a string in the source JSON files. Which TWO strategies can be used to manage this schema conflict?
Hard78A data engineer runs a nightly Databricks job that ingests data into a Delta table using multiple concurrent write streams. The engineer notices that some write transactions are failing with a ConcurrentAppendException. The job writes to the same partition of the table from different tasks. Which action should the engineer take to resolve this issue?
Medium79A data engineer is troubleshooting a slow-running Spark job on Databricks. Which TWO metrics in the Spark UI are most useful for identifying data skew?
Medium80Refer to the exhibit. A CI/CD pipeline running a Databricks CLI command fails with the error shown. What is the most likely cause?
Medium81A data engineer is designing an access control model in Unity Catalog for a new catalog `finance`. The team wants to follow the principle of least privilege while still enabling collaboration. Which TWO of the following practices best align with Unity Catalog's privilege model? (Choose two.)
Medium82What is the primary function of a Delta Lake 'Vacuum' operation?
Easy83A data engineer is monitoring a Databricks job and notices that the job's duration has gradually increased over the past week. The job reads a large Delta table, performs aggregations, and writes results to another Delta table. The engineer wants to identify the stage that is taking the most time. Which Spark UI tab should the engineer use to quickly identify the slowest stage?
Easy84A data engineer is investigating why a Databricks job that writes to a Delta table is experiencing performance degradation over time. The job performs frequent small appends. Which TWO actions should the engineer take to improve write performance? (Choose two.)
Hard85A data engineer needs to share a table in Unity Catalog with a partner organization using a different Databricks account. Which feature should be used to provide secure access without moving the data?
Medium86An engineer needs to identify the root cause of a job failure. Which THREE of the following are valid locations or methods to investigate the logs?
Medium87A data platform team is migrating their deployment process to Databricks Asset Bundles (DABs). They already have a Python wheel task defined in a Databricks Job and a set of notebooks in a Git repository. They want the bundle deployment to be repeatable across development, staging, and production targets with different cluster sizes. Which approach should they take to parameterize the target-specific cluster configuration?
Medium88A data engineer wants to use 'Expectations' in Delta Live Tables to monitor data quality. What happens if a record violates an expectation defined with the 'fail' constraint?
Medium89A data engineer has a Lakeflow Job with three tasks: bronze_ingest, silver_transform, and gold_aggregate. The silver_transform task must run only if bronze_ingest succeeds, and gold_aggregate must run only if silver_transform succeeds. The engineer also wants gold_aggregate to run even if silver_transform fails, so that partial results can be published. Which configuration should the engineer apply to gold_aggregate?
Hard90A data engineer is working on a Delta Lake table that has accumulated millions of small files due to frequent streaming updates. This fragmentation has significantly degraded query performance. Which operation should the engineer execute to optimize file layout without altering table data?
Medium91Which of the following is the primary benefit of using a 'Job Cluster' rather than an 'All-Purpose Cluster' for running scheduled data pipelines?
Easy92A data engineer has a Unity Catalog table `prod.sales.orders` that contains a column `customer_email`. They need to allow analysts in the `marketing` group to query the table but only see a masked version of `customer_email` (e.g., `a***@example.com`). The masking logic is implemented as a SQL user-defined function `prod.security.mask_email`. Which statement should the engineer execute to apply the mask?
Medium93A data engineering team stores all production notebooks and job definitions in a Git repository. They want every merge to the `main` branch to automatically deploy the updated notebooks to the production Databricks workspace without any manual copy/paste. Which approach should they implement?
Medium94A data engineer needs to troubleshoot a job that is failing during the 'shuffle' phase. Which Spark UI tab should the engineer examine to analyze the shuffle partitions and identify potential imbalances?
Medium95A data engineer maintains a Delta table named inventory.products with columns product_id, category, price, and updated_at. The engineer needs to create a new table that contains one row per category with the average price and the most recently updated product_id in that category. The query must be efficient and use only standard Databricks SQL. Which statement should the engineer run?
Medium96A data engineer is creating a Silver table in a Delta Live Tables pipeline. The pipeline must continuously ingest new files from a cloud storage location as they arrive, and the engineer wants to avoid reprocessing files that were already ingested. Which approach should be used to read the source data?
Easy97An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?
Medium98A data engineer is building a Gold aggregate table that summarizes daily sales by product category. The Silver source is a streaming Delta table that receives late-arriving events up to 48 hours old. The engineer needs the Gold table to always reflect the most accurate aggregates, including corrections for late data, without full recomputation. Which approach is most appropriate?
Hard99Refer to the exhibit. An engineer notices that queries filtering by 'customer_id' are running slowly on the 'orders' table. Based on the exhibit, what is the most appropriate action to resolve this?
Medium100A data engineer is using Databricks Jobs to run a nightly ETL pipeline. The job occasionally fails due to a transient network error when writing to an external database. The engineer wants to automatically retry the job a few times before marking it as failed. What is the most efficient way to configure this in Databricks?
Medium101Refer to the exhibit. If 'task1' fails due to a timeout, what happens to 'task2'?
Hard102Which component of Databricks CI/CD is responsible for executing automated tests on code before it is merged into the main branch?
Easy103Which command is used to query the history of a Delta table to perform time travel?
Medium104A data engineer manages a Unity Catalog table `sales.raw.transactions` that contains a column `credit_card_number`. The security team requires that users in the `auditors` group can see the full credit card number, while all other users who have SELECT privileges on the table should see only the last four digits. The engineer decides to use a column mask. Which SQL statement should the engineer execute to meet this requirement?
Hard105A data engineer is designing a Delta Lake Bronze-to-Silver pipeline in Databricks and needs to ensure that downstream consumers receive high-quality data. Which TWO data quality enforcement mechanisms are natively supported in Delta Live Tables using expectations?
Hard106A data engineer needs to ingest a large CSV file from cloud storage into a Delta table using Databricks SQL. The engineer wants to perform a one-time load and ensure that the operation is atomic. Which command should be used?
Easy107Which of the following is the best practice for managing service principals in a Databricks workspace?
Medium108Refer to the exhibit. A data engineer is configuring a streaming pipeline. Which outcome will these specific configurations have on the target table's performance?
Medium109A team is implementing a CI/CD process for their Delta Live Tables (DLT) pipelines. Which THREE of the following practices are recommended to ensure reliable deployment?
Hard110Refer to the exhibit. A CI/CD pipeline fails with the provided error. What is the most likely cause?
Hard111A team wants to automate deployment of Databricks jobs from a Git repository using a CI/CD pipeline. They need a tool that reads a declarative project definition and creates or updates jobs, pipelines, and notebooks in a target workspace. Which Databricks capability should they use?
Easy112A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?
Hard113An engineer notices that a specific notebook job is consistently taking longer to start. They observe high 'initialization' times in the job logs. Which action should the engineer take to improve startup time?
Medium114A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)
Hard115A data engineer is monitoring a Databricks job and notices that the job's tasks are spending a significant amount of time in garbage collection (GC). The job processes large amounts of data with many small objects. Which action should the engineer take to reduce GC overhead?
Medium116A data engineer is using Databricks SQL to analyze data stored in a Delta table. The engineer wants to optimize query performance by leveraging Delta Lake features. Which TWO actions should the engineer take to improve query performance on the Delta table? (Choose two.)
Medium117A data engineer runs a Structured Streaming job that writes to a Delta table. The job processes data from a Kafka topic and uses a 10-minute watermark. After a few hours, the engineer notices that the streaming query's input rate is steady, but the processing rate has dropped significantly, and the batch duration has increased from 5 seconds to over 2 minutes. The job is running on a cluster with autoscaling enabled. Which action should the engineer take FIRST to diagnose the performance degradation?
Medium118A data engineer needs to share a table in Unity Catalog with external partners who do not have access to the Databricks workspace. Which feature should the data engineer configure?
Medium119Which TWO of the following statements are correct regarding the use of Databricks Repos for CI/CD?
Medium120Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?
Easy121A data engineer is working with a Delta table that contains a column named raw_data of type STRING, which holds JSON strings. The engineer needs to extract specific fields from this JSON and store them as separate columns in a new Delta table. Which approach is most efficient and maintains data quality?
Easy122A data engineer is optimizing a Delta table that suffers from slow read performance due to small file sizes. Which command should the engineer execute to consolidate these small files into larger, more efficient files without altering the underlying table data?
Medium123A data engineer needs to monitor the costs associated with specific projects running on a shared Databricks workspace. Which feature should the engineer use to attribute these costs accurately?
Easy124You are tasked with handling late-arriving data in a streaming pipeline that performs windowed aggregations. Which approach ensures that the output remains accurate while balancing memory usage?
Medium125A data engineer is using Databricks Auto Loader to stream CSV files from an ADLS Gen2 container into a Delta table. The source directory contains a mix of files, but only files with the prefix 'sales_' should be ingested. The engineer wants Auto Loader to ignore all other files without moving or deleting them. Which Auto Loader option should the engineer configure to achieve this?
Medium126An analytics team needs to frequently query a large Delta table by a high-cardinality customer_id column and a date column. To optimize query performance and reduce data scanning during filtering, how should the data engineer structure the table layout?
Medium127A data engineer has a Unity Catalog table `prod.sales.orders` containing a column `customer_email`. Company policy requires that users in the `analyst_group` see only the domain part of the email (e.g., `***@example.com`), while members of `pii_admin_group` must see the full email. The engineer wants to enforce this at query time without creating separate views. Which approach should the engineer use?
Medium128A data engineering team stores its Databricks notebooks and Python files in a Git repository. They want to avoid manually copying files into the workspace and ensure that the production workspace always runs the exact code version that passed tests. Which approach should they use?
Medium129A data engineer is optimizing a Databricks job that reads from a large Delta table and performs a join with a smaller table. The job is experiencing performance issues due to shuffling. The engineer wants to reduce the amount of data shuffled during the join. Which technique should the engineer use?
Hard130A data engineering team stores notebooks in a Git repository and wants automated deployments to a Databricks workspace. They configure a GitHub Actions workflow that runs a Databricks CLI command to deploy Databricks Asset Bundles. The workflow authenticates using a service principal OAuth token stored in GitHub Secrets. After the first successful run, subsequent runs fail with an authentication error. The token was created with a 1-hour lifetime. What should the team do to ensure the workflow can authenticate reliably on every run?
Medium131A data engineer is configuring a storage credential in Unity Catalog to access an AWS S3 bucket. The engineer creates an IAM role with the necessary permissions and sets up the storage credential using the role ARN. What additional step is required to allow Databricks to assume the role?
Medium132A data engineer stores a Delta table in Unity Catalog at the managed location of the schema 'sales'. The table is later dropped using DROP TABLE. What happens to the underlying data files?
Medium133A data engineer is creating a Lakeflow Job that must run a notebook every weekday at 06:00 in the company's local time zone, which is America/New_York. The engineer configures a schedule trigger but the job runs at the wrong time. Which setting should the engineer verify first?
Easy134A team is preparing to optimize their Databricks data transformation pipeline. Which THREE of the following actions are considered best practices for optimizing Delta Lake performance?
Medium135A team is using Databricks Asset Bundles (DABs) to manage their CI/CD pipeline. They want to run unit tests on their Python code before deploying the bundle. Where should the tests be executed in the pipeline?
Medium136A data engineer has a Lakeflow Job that runs daily. They want to receive an email only when the job fails, not on every run. Which notification configuration should they set?
Easy137A data engineer is designing a pipeline on Databricks that requires ACID transactions and schema enforcement for streaming data. Which storage abstraction should they use to ensure data reliability and support time travel?
Medium138A data engineer wants to pass a file path from Task A to Task B in a Lakeflow Job. Task A is a notebook that computes the path, and Task B is a notebook that reads from that path. Which mechanism should the engineer use to share the value between tasks?
Medium139Which Databricks feature should be used to securely share data with external organizations without duplicating the data?
Medium140A data engineer has a Lakeflow Job with a linear dependency chain: Task A, then Task B, then Task C. Task B sometimes fails due to transient errors. The engineer wants Task C to run only if Task B succeeds, but also wants Task B to be retried automatically before considering the job failed. Which configuration should they use?
Hard141Which object type in Databricks Unity Catalog acts as the top-level container for organizing schemas and tables, providing a unified namespace for data assets?
Easy142Why is it important to use Service Principals instead of Personal Access Tokens (PATs) for CI/CD automation in Databricks?
Easy143Which of the following is a recommended strategy for managing library dependencies in a CI/CD pipeline for Databricks?
Medium144A data engineer needs to grant a group `data_consumers` the ability to query a view `prod.reporting.sales_summary` that is defined on top of tables in the same catalog. The group currently has no privileges on the underlying tables. What is the minimum set of privileges the engineer must grant to `data_consumers` so they can query the view successfully?
Medium145A data engineer is analyzing a Spark job that is failing with 'Out of Memory' (OOM) errors. Which configuration parameter should be tuned to increase the amount of memory allocated to the execution of joins and aggregations?
Medium146Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?
Easy147Which of the following is the primary benefit of using 'Auto Loader' (cloudFiles) for data ingestion in Databricks compared to standard batch processing?
Easy148Refer to the exhibit. When is this job scheduled to run?
Hard149An organization requires that all data access logs across multiple Databricks workspaces be captured and sent to a centralized security information and event management (SIEM) system. Which Databricks feature should be configured to capture these audit events?
Medium150A data engineer is configuring a Unity Catalog storage credential to access an AWS S3 bucket. The S3 bucket policy grants access to an IAM role. The data engineer creates a storage credential with that IAM role's ARN. However, when attempting to create an external location using this storage credential, the operation fails with an error indicating insufficient permissions. The data engineer verifies that the IAM role has the correct S3 permissions. What is the most likely cause of the failure?
Hard151A data engineer is building a Silver table in Delta Lake from a Bronze table that contains raw JSON events. The engineer needs to flatten a nested struct column named 'device' with fields 'type' and 'os', and also extract a field from an array of structs named 'events'. The goal is to produce a clean, denormalized Silver table. Which PySpark operation should the engineer use to achieve this transformation efficiently?
Medium152A data engineering team is implementing a CI/CD pipeline for Databricks notebooks and jobs using Databricks Asset Bundles. They want to ensure deployments are reproducible and that production changes are traceable. Which TWO of the following practices should they follow? (Choose two.)
Hard153Refer to the exhibit. A data engineer is using this command to reload data into a Delta table after a schema correction. What is the primary effect of setting the 'force' option to 'true' in this context?
Medium154A data engineer is setting up a Lakeflow Job that runs a notebook task. The engineer needs the task to always execute even if the upstream task in the workflow fails. Which configuration should be applied to the dependent task's condition?
Medium155A data engineer is designing a pipeline and needs to ensure that data remains consistent during concurrent read and write operations. Which Databricks feature provides the mechanism to track and validate these operations?
Medium156A data engineering team must implement dynamic row-level filtering on a customer analytics table in Unity Catalog so that regional analysts only view records corresponding to their assigned territory. Which TWO steps are required to achieve this using Unity Catalog features? (Choose 2)
Hard157What is the primary goal of implementing 'environment parity' in a Databricks CI/CD pipeline?
Medium158A data engineer wants to monitor the health and performance of Databricks Jobs over time. Which feature should they use to visualize trends, such as job success rates and average execution times, across multiple runs?
Medium159A data engineering team needs to ingest streaming data from Kafka into a Delta table while maintaining exactly-once processing guarantees and low latency. Which Databricks Intelligence Platform feature should they utilize to build this streaming pipeline declaratively?
Medium160Refer to the exhibit. The streaming pipeline is experiencing memory issues because the state store size increases continuously. What is the most effective way to address this while maintaining aggregation accuracy?
Hard161A data engineer needs to ensure that all queries against a Unity Catalog table are logged for audit purposes. The logs must include the user identity, the query text, and the timestamp. Which Databricks feature should be enabled to capture this information?
Easy162A data engineer is migrating legacy batch jobs to Delta Live Tables (DLT). They want to optimize performance for a complex join operation between two large tables. Which TWO strategies should they implement to improve the join efficiency?
Medium163A team is using the Databricks CLI in their CI/CD pipeline to deploy jobs and notebooks. They need the pipeline to authenticate to a production workspace without embedding a personal user's credentials. Which authentication method should they configure for the CLI?
Easy164Which TWO of the following statements accurately describe the functionality of Unity Catalog within the Databricks Intelligence Platform?
Medium165A data engineer has configured a storage credential in Unity Catalog to access an AWS S3 bucket. The credential uses an IAM role with a trust policy that allows Databricks to assume it. The engineer now needs to ensure that only a specific set of users can create external tables pointing to that S3 bucket. Which Unity Catalog object should be used to control this access?
Hard166When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?
Easy167A data engineer is investigating a Databricks job that failed overnight. The job's status in the Jobs UI shows 'Failed', and the engineer needs to view the error message and stack trace to determine the cause. Where should the engineer look to find the detailed error information for the failed run?
Easy168Which TWO of the following are primary benefits of using Delta Live Tables (DLT) for data pipeline development?
Medium169A data engineering team wants to implement Git integration for their Databricks notebooks. Which workflow is considered the best practice for CI/CD in Databricks Repos?
Medium170A data engineer is setting up a CI/CD pipeline that runs unit tests on transformation logic before deploying notebooks to a production Databricks workspace. The tests must run quickly and not depend on a live Databricks cluster or external data sources. Which approach best meets these requirements?
Medium171A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?
Medium172Which THREE of the following represent core security principles enforced by Unity Catalog?
Medium173You are building a pipeline where a Bronze table contains JSON data with a nested 'user_info' struct. You need to promote this to a Silver table where 'user_id' is a top-level column. Which approach is the most efficient for this transformation?
Medium174A data engineer is building a Delta Live Tables pipeline that ingests streaming data from a Kafka topic into a bronze table, then applies a series of transformations to produce a silver table. The engineer notices that the pipeline is reprocessing all data from the beginning of the Kafka topic on each run, causing high latency. The Kafka topic has a retention period of 7 days, and the pipeline is configured to use the default settings. What is the most likely cause of this behavior?
Hard175A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?
Medium176A data engineer is asked to ensure that all queries against a Unity Catalog table are recorded for auditing, including the identity of the user and the query text. Which Databricks feature should the engineer enable to capture this information?
Easy177Which THREE of the following are primary components of the Databricks Lakehouse architecture?
Medium178Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?
Medium179A data engineering manager wants to ensure that all production code in Databricks is fully audited and versioned. Which TWO of the following steps are required?
Hard180A data engineer is using Databricks Auto Loader to ingest JSON files from a directory. The stream is configured with `cloudFiles.schemaLocation` set to a specific path. The engineer notices that the stream fails when a new file contains an additional column. What is the most likely reason for the failure?
Hard181A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?
Medium182A data engineer configures a Lakeflow Job to run a notebook task on a job cluster. The notebook reads a parameter named run_date using the widget API. During a manual run, the engineer wants to supply a specific date without editing the notebook. The job also runs on a nightly schedule where the date should default to the current day. Which approach correctly supplies the parameter for both the manual and scheduled runs?
Hard183A data engineer needs to share a Delta table managed by Unity Catalog with external partners who do not have access to the Databricks workspace. Which feature should be used to securely grant read-only access to this table without replicating the data?
Medium184A data engineer is designing a Delta Lake pipeline that processes streaming sales transactions. The schema evolves frequently, and the pipeline must handle these changes without manual intervention. Which feature should the engineer enable to support automatic schema updates while preventing data corruption?
Medium185What is the primary role of a 'Metastore Admin' in a Databricks Unity Catalog environment?
Medium186A team uses Databricks Asset Bundles to define a job that must exist in both a staging and a production workspace with different cluster sizes. They want a single bundle definition that deploys correctly to both targets. Which configuration should they use?
Medium187A data engineer is setting up a Databricks job that runs a notebook on a schedule. The job must process data from a source that is updated daily and write results to a Delta table. The engineer wants to ensure that if the job fails, it automatically retries up to three times. Which feature should the engineer configure in the job settings to achieve this?
Medium188A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?
Medium189Your team is migrating a manual Databricks job to a CI/CD pipeline. You need to ensure the job configuration is version-controlled and deployed programmatically. Which approach aligns with Databricks best practices?
Medium190Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?
Medium191A data engineer is designing an ingestion pipeline using Databricks Auto Loader to process JSON files from an S3 bucket. The pipeline must handle schema evolution and ensure that data is ingested exactly once. Which two features of Auto Loader support these requirements? (Choose two.)
Hard192A data engineer wants to move data from a 'Bronze' table to a 'Silver' table while performing data cleaning. They want to ensure that this process is only executed once per data batch. Which approach is best for this requirement?
Medium193Refer to the exhibit. A data engineer receives this error when collecting data from a large transformation back to the driver node. Which approach should be used to fix this issue?
Hard194Which TWO of the following statements accurately describe the behavior of Delta Lake table constraints and enforcement?
Hard195A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing mode to file notification mode. What is the primary reason for making this change?
Hard196What must be configured to allow Databricks to access cloud storage on behalf of a user without the user needing to provide their own cloud credentials?
Medium197A data engineer needs to join two massive datasets. One dataset is very small (10MB), and the other is very large (1TB). To ensure the join operation is performed as efficiently as possible, which join strategy should be enforced?
Medium198Refer to the exhibit. A Databricks administrator is using the Unity Catalog JSON policy to manage access. If the 'data_scientist_1' user attempts to execute an 'UPDATE' command on the 'orders' table, what will be the result?
Hard199A data engineer has a Lakeflow Job with a notebook task that occasionally fails due to transient network errors when reading from an external REST API. The engineer wants the task to automatically retry up to three times, but only for this specific task, without affecting other tasks in the job. What should the engineer do?
Medium200A data engineer is working in a Databricks workspace where Unity Catalog is enabled. They need to run a SQL query that reads from the table sales in the catalog prod and schema marketing. Which fully qualified name should they use?
Easy201Which capability is provided by Databricks' integration with MLflow?
Easy202While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?
Medium203A data engineer is using Databricks Asset Bundles to deploy a data pipeline that includes a job and a notebook. The engineer wants to ensure that the deployment is idempotent and can be rolled back if needed. Which TWO of the following statements accurately describe the benefits of using Databricks Asset Bundles for this scenario? (Choose two.)
Medium204Refer to the exhibit. An organization needs to ensure that members of the 'analyst_group' can only view data where the region column equals 'US'. Based on the provided configuration, what is the best approach to implement this in Unity Catalog?
Medium205Refer to the exhibit. An engineer applies these configurations to a cluster. What is the primary benefit of enabling the Databricks IO Cache for a workload that involves repeatedly reading the same Delta tables?
Hard206Refer to the exhibit. A data engineer is attempting to run a VACUUM command on a table, but the command fails with an error indicating that the retention period is too short. Given the configuration, what is the most appropriate action the engineer should take to safely remove files older than 7 days?
Hard207A data engineer is optimizing a Databricks job that processes a large dataset. The job performs a join between a large Delta table and a small dimension table, then writes the result to a Delta table. The engineer notices that the join is causing a large shuffle and wants to reduce shuffle overhead. Which two actions should the engineer take to improve performance? (Choose two.)
Medium208A data engineer is building a Lakeflow Job that must run a sequence of tasks across different compute types. The ingest task must run on a job cluster with a specific Spark configuration, the transform task must run as a Delta Live Tables pipeline, and the report task must run on a separate SQL warehouse. Which TWO statements about task-level compute configuration in Lakeflow Jobs are correct? (Choose two.)
Hard209A team is using Databricks Asset Bundles (DABs) to deploy a job to multiple environments (dev, staging, prod). They need to ensure that the job uses different cluster sizes and schedules per environment. Which DABs feature should they use?
Hard210A data engineer is using Databricks Auto Loader to ingest JSON files from an Azure Data Lake Storage Gen2 container into a Delta table. The engineer notices that the ingestion is slow and wants to optimize file discovery. The directory contains millions of files, and new files are added frequently. Which Auto Loader option should be used to improve file discovery performance?
Hard211A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.json. The engineer wants to automatically add a column 'date' to the ingested data based on the file path. Which Auto Loader feature should be used?
Medium212A data engineer is building a Lakeflow Job with a task that runs a SQL notebook. The task must run only on weekdays and must be completed before 9 AM. The engineer wants to configure the schedule to meet these requirements. What should the engineer do?
Medium213A data engineer needs to optimize the layout of a massive Delta Lake table that suffers from poor query performance due to a large number of small files and unsorted data records. Which TWO operations should the engineer execute to resolve these performance bottlenecks?
Medium214A data engineer is designing a pipeline on Databricks to process streaming data. Which architectural component acts as the unified storage layer, allowing both batch and streaming workloads to access the same underlying data files in a data lake?
Medium215Which feature of the Databricks Intelligence Platform allows users to manage fine-grained access control across workspaces for tables, files, and machine learning models?
Easy216A data engineer is troubleshooting a Databricks job that intermittently fails with a `SparkException: Job aborted due to stage failure: Task not serializable`. The job reads from a Parquet file, performs a transformation using a custom function defined in a Python class, and writes to a Delta table. The engineer suspects that the custom function is causing the issue. Which action should the engineer take to resolve the serialization error?
Hard217Refer to the exhibit. A data engineer encounters this error while trying to list tables in a schema. Based on the error, what must the engineer do to resolve it?
Hard218Which THREE of the following are benefits of using Unity Catalog over the legacy Hive Metastore?
Medium219A data engineer has a Unity Catalog table named `sales.raw.orders` that contains a column `credit_card` with sensitive data. The security team requires that users in the `analyst` group see only the last four digits of the credit card number when querying this table, while all other users with appropriate privileges see the full value. The data engineer wants to implement this with minimal disruption to existing queries. Which approach should the data engineer take?
Medium220A data engineer is investigating why a Databricks job that reads from a Delta table is slow. The job performs a simple SELECT with a filter on a partition column. The engineer suspects that the table has many small files. Which Spark UI tab should be examined to confirm the number of files read?
Easy221A data engineer is configuring a CI/CD pipeline for Databricks notebooks using GitHub Actions. They need to authenticate to Databricks to deploy notebooks. Which authentication method is recommended for production CI/CD pipelines?
Medium222What is the primary function of the 'Retries' setting in a Databricks Job task?
Medium223A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?
Medium224A team stores Databricks notebooks and job definitions in a Git repository and wants every merge to the main branch to automatically deploy to production. Which combination of practices should the pipeline implement to achieve this safely?
Medium225Which TWO of the following statements correctly describe the behavior of the Delta Lake 'MERGE' operation when handling schema evolution?
Hard226When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?
Medium227A data engineer is designing a workflow that requires running a series of tasks with dependencies. The workflow must be able to retry failed tasks automatically and send notifications on failure. The engineer also needs to ensure that the workflow can be triggered on a schedule and via an API. Which Databricks feature should the engineer use?
Hard228A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?
Hard229A data engineer is working with a large Delta table and notices that queries filtering by 'region_id' are performing slowly. The table is currently partitioned by 'date'. Which strategy should the engineer use to optimize query performance for 'region_id' filtering without increasing the number of partitions?
Medium230A data engineer needs to configure a Databricks Job to orchestrate a data pipeline that includes a Python task, a SQL task, and a notebook task. The pipeline requires passing a dynamic run identifier from the Python task to the subsequent SQL and notebook tasks. Which mechanism should the data engineer use to achieve this task-to-task dependency parameter passing?
Medium231Refer to the exhibit. The merge operation is failing in your pipeline. What is the root cause of this error, and how should it be resolved?
Hard232When utilizing Databricks Asset Bundles, how should secrets (e.g., API keys, database credentials) be handled to ensure security during CI/CD?
Hard233A data engineer is configuring Unity Catalog row-level security on a table `prod.finance.transactions`. They want to ensure that users in the `us_team` group can only see rows where `region = 'US'`, while users in the `eu_team` group can only see rows where `region = 'EU'`. They create a row filter function `prod.security.region_filter`. Which TWO statements accurately describe how to implement and manage this row-level security? (Choose two.)
Hard234Refer to the exhibit. A data engineer is configuring an IAM policy for a Databricks storage credential. Which of the following is the most significant security risk in this configuration?
Hard235Why should developers avoid using 'notebook' references in production pipelines that point to the 'Shared' folder for shared development work?
Medium236A data engineer working in a Databricks workspace needs to create a new notebook that will be shared with teammates in the same workspace. They want the notebook to be organized in a folder named 'team_project' and be visible to all workspace users. Which action should the data engineer take to accomplish this in the Databricks workspace?
Easy237When integrating Databricks with a CI/CD tool like GitHub Actions, how should you securely manage the authentication token used for deployments?
Easy238Which of the following best describes the purpose of 'Credential Passthrough' in Databricks?
Hard239A data engineer is using Auto Loader to ingest JSON files from a cloud storage directory into a Delta table. The directory receives files continuously, and the engineer wants to ensure that the ingestion process can handle schema drift where new columns are added to the JSON files over time. The engineer also wants to minimize the need to reprocess all data when the schema changes. Which configuration should the engineer use?
Hard240A data engineer manages a Lakeflow Job that runs a long-running notebook task on a job cluster. The task occasionally fails due to transient cloud storage errors, and the engineer wants the task to retry automatically without failing the entire job on the first attempt. The engineer also wants to be alerted only if all retries are exhausted. Which configuration should the engineer apply?
Medium241A data engineer is implementing a Type 2 slowly changing dimension in Delta Lake for a customers table. The table has columns customer_id, name, address, effective_date, end_date, and is_current. When a customer's address changes, the engineer wants to expire the existing current row and insert a new current row in a single atomic operation. Which Delta Lake feature should the engineer use?
Medium242Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?
Medium243A team is transitioning to the Databricks Intelligence Platform. Which TWO actions are required to successfully register a table in Unity Catalog using the three-level namespace?
Hard244A CI/CD pipeline must run unit tests on Python transformation code before deploying a Databricks job. The tests should execute quickly without starting a cluster and should validate the transformation logic in isolation. Which approach best meets these requirements?
Hard245A data engineer needs to share a Delta table with an external partner who does not have a Databricks account. The engineer wants to provide read-only access to the table for a limited time, ensuring the partner cannot access any other data. Which Databricks feature should the engineer use?
Medium246A data engineer maintains a Delta table where each row represents a customer record, and updates arrive continuously as change data capture events. The engineer needs to apply inserts, updates, and deletes from a staging table into the target table in a single atomic operation, matching records on customer_id. Which Delta Lake operation should be used?
Hard247A data engineer needs to ingest data from a legacy SQL Server database into a Delta Lake bronze table. The ingestion must be performant and support parallel reads from the source table. What is the best practice for configuring the JDBC connection in this scenario?
Medium248When designing an ingestion strategy for a Delta Lake architecture, which TWO advantages does Delta Lake provide over traditional Parquet tables for incoming data?
Medium249Which security feature in Databricks allows administrators to mask sensitive data, such as email addresses or social security numbers, in query results based on user-defined functions?
Easy250A data engineer needs to run a Databricks notebook that processes data stored in an external ADLS Gen2 location. The notebook must access the data securely without embedding credentials in the notebook code. The engineer has already configured a Unity Catalog external location with a storage credential. Which method should the engineer use to read the data?
Easy251What is the primary function of an 'Access Connector' for Azure Databricks when using Unity Catalog?
Easy252Refer to the exhibit. A data engineer is attempting to ingest JSON data into a Delta table. Based on the error log, what is the most appropriate transformation step to implement before loading this data into the production table?
Medium253A junior data engineer writes a PySpark transformation that reads a Parquet dataset, filters out inactive users, and appends the resulting DataFrame to an existing Delta Lake table. However, the data engineer notices duplicate records appearing in the target table after multiple runs. Which technique should be implemented to ensure idempotency?
Easy254A team is using Databricks Asset Bundles to manage a job that writes to a Unity Catalog table. The bundle is deployed to a staging workspace for testing and then to a production workspace. The team wants the job to use different catalog and schema names in each environment without duplicating the entire bundle. Which approach should they use?
Hard255Which TWO of the following capabilities are native to the Databricks Unity Catalog?
Medium256Which TWO of the following are benefits of using Unity Catalog over the legacy Hive Metastore?
Medium257A data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?
Medium258A data engineer is configuring Unity Catalog to allow a service principal to run a Databricks job that writes to a Delta table in an external location. The service principal has been granted USAGE on the catalog and schema, and MODIFY on the table. However, the job fails with an error indicating that the service principal cannot access the external location. What is the most likely missing privilege?
Medium259A data engineer is designing a Delta Lake table that will be used for both batch and streaming reads. The table must support upserts from a streaming source and maintain ACID transactions. The engineer wants to ensure that the table can be efficiently queried by downstream consumers using SQL while minimizing storage costs. Which TWO actions should the engineer take to meet these requirements? (Choose two.)
Hard260A data engineer needs to automate the ingestion and transformation of data within Databricks with minimal manual intervention. Which feature is most appropriate for orchestrating these data workflows?
Medium261A data engineer has a Delta table named silver_events with columns event_id (string), event_ts (timestamp), and payload (string). The table is partitioned by event_date (derived from event_ts). The engineer needs to update the payload column for all events that occurred on '2024-06-01' based on a mapping table named updates (event_id, new_payload). Which PySpark operation should be used to perform this update efficiently while preserving Delta Lake ACID guarantees?
Medium262When a Data Engineer uses a 'Repair and Rerun' functionality on a failed Databricks Job, what happens?
Medium263A team wants to ensure that code changes to their Databricks notebooks are reviewed before being deployed to production. They use a Git repository and Databricks Repos. Which practice should they implement in their CI/CD process?
Easy264Which identity management approach is mandatory for the implementation of Unity Catalog within a Databricks environment?
Medium265A CI/CD pipeline runs unit tests against transformation logic before deploying to production. The tests must run on a Databricks cluster and produce a pass/fail result that fails the pipeline when assertions do not hold. Which implementation best fits this requirement?
Hard266A data engineer needs to perform an upsert on a target Delta table using a source DataFrame. Which operation provides the most robust mechanism to handle duplicates and updates in a single pass?
Medium267A data engineer needs to create a Delta table in Unity Catalog that will be used by multiple teams. The engineer wants to ensure that the table is governed by Unity Catalog and that access can be controlled using SQL GRANT statements. What is the correct way to create the table?
Easy268Refer to the exhibit. A data engineer is deploying an automated job. Based on the provided configuration, what is the primary benefit of using the 'autoscale' attribute in this cluster definition?
Medium269When a data engineer needs to automate a recurring ETL job, which Databricks tool is most appropriate for orchestrating tasks and handling dependencies?
Medium270A data engineer needs to run a Databricks notebook on a schedule every day at 8:00 AM UTC. The notebook performs data transformations and writes results to a Delta table. The engineer wants to ensure the job runs reliably and can be monitored. Which Databricks feature should the engineer use to schedule and monitor the notebook?
Easy271A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?
Medium272A data engineer needs to ensure that only members of the 'finance' group can view a specific column containing credit card numbers in a Unity Catalog table. Which feature should be used?
Easy273A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?
Hard274Which Databricks compute resource is specifically optimized for running BI dashboards and SQL queries?
Easy275A data engineer maintains a Lakeflow Job with a scheduled trigger set to run every day at 08:00. The job's source table is refreshed by an upstream process that sometimes finishes later than expected, causing the job to process stale data. The engineer wants the job to start only after the upstream refresh completes, regardless of the clock time, while still preserving the existing 08:00 schedule as a fallback. Which trigger configuration should the engineer implement?
Medium276A data engineer needs to join a small 'lookup' table with a large 'transactions' table in a Spark job. Which transformation strategy will provide the best performance in a cluster environment?
MediumOther domains
All Databricks-DE-Assoc exam domains
Frequently asked questions
- What does the scenario questions domain cover on the Databricks-DE-Assoc exam?
- scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
- How many questions are in this domain?
- This page lists all 276 scenario questions questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only scenario questions questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.