Sample questions
Databricks Certified Data Engineer Associate practice questions
A data engineer is using Auto Loader to ingest data from a directory that receives thousands of files every hour. They are considering switching from the default directory listing…
A data engineer has a Lakeflow Job with a notebook task that occasionally fails due to transient network errors when reading from an external REST API. The engineer wants the task…
While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture the…
Refer to the exhibit. An organization needs to ensure that members of the 'analyst_group' can only view data where the region column equals 'US'. Based on the provided configuratio…
A data engineer is designing a secure architecture using the Databricks Intelligence Platform. Which TWO of the following statements accurately describe the role and capabilities o…
A team is using Databricks Asset Bundles (DABs) to deploy a job to multiple environments (dev, staging, prod). They need to ensure that the job uses different cluster sizes and sch…
A data engineer is using Auto Loader to ingest files from an S3 bucket into a Delta table. The files are partitioned by date in the path, e.g., s3://bucket/data/2023-01-01/file1.js…
Which feature of the Databricks Intelligence Platform allows users to manage fine-grained access control across workspaces for tables, files, and machine learning models?
An organization requires that all data access logs across multiple Databricks workspaces be captured and sent to a centralized security information and event management (SIEM) syst…
What is the primary function of the 'Retries' setting in a Databricks Job task?
A team stores Databricks notebooks and job definitions in a Git repository and wants every merge to the main branch to automatically deploy to production. Which combination of prac…
When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?
Troubleshooting, Monitoring, and OptimizationmediumSee the answer and why each option is right or wrong →A data engineer is designing a workflow that requires running a series of tasks with dependencies. The workflow must be able to retry failed tasks automatically and send notificati…
Which TWO of the following capabilities are native to the Databricks Unity Catalog?
A data engineer is configuring Unity Catalog row-level security on a table `prod.finance.transactions`. They want to ensure that users in the `us_team` group can only see rows wher…
A data engineer is working with a large Delta table and notices that queries filtering by 'region_id' are performing slowly. The table is currently partitioned by 'date'. Which str…
Which feature in Unity Catalog is primarily used to track data movement and transformation history for compliance and auditing?
When integrating Databricks with a CI/CD tool like GitHub Actions, how should you securely manage the authentication token used for deployments?
A data engineer wants to run an incremental ingestion job every six hours. They want to ensure that each run processes all available data and then shuts down the cluster to save co…
Which of the following best describes the purpose of 'Credential Passthrough' in Databricks?
A data engineer is implementing a Type 2 slowly changing dimension in Delta Lake for a customers table. The table has columns customer_id, name, address, effective_date, end_date,…
Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?
Troubleshooting, Monitoring, and OptimizationmediumSee the answer and why each option is right or wrong →A team is transitioning to the Databricks Intelligence Platform. Which TWO actions are required to successfully register a table in Unity Catalog using the three-level namespace?
A team wants to ensure that code changes to their Databricks notebooks are reviewed before being deployed to production. They use a Git repository and Databricks Repos. Which pract…