Databricks · Free Practice Questions · Last reviewed May 2026
42real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A data engineer is designing a Delta Lake Bronze-to-Silver pipeline in Databricks and needs to ensure that downstream consumers receive high-quality data. Which TWO data quality enforcement mechanisms are natively supported in Delta Live Tables using expectations?
CONSTRAINT expectation_name EXPECT (column_name IS NOT NULL) ON VIOLATION DROP ROW
This declarative expectation syntax successfully instructs Delta Live Tables to validate the specified condition and automatically drop any incoming records that violate the constraint, ensuring only clean data persists in the target table.
CONSTRAINT expectation_name EXPECT (column_name IS NOT NULL) ON VIOLATION FAIL UPDATE
This native expectation clause correctly tells the Delta Live Table pipeline to immediately halt execution and fail the update if any incoming record fails the validation rule, preventing corrupted or incomplete datasets from propagating.
ASSERT expectation_name ON VIOLATION RETRY
FILTER expectation_name WHERE column_name IS NOT NULL
VALIDATE expectation_name ON ERROR IGNORE
A data engineer is designing a Bronze-to-Silver transformation pipeline using Delta Lake. They need to ensure that the Silver table contains only records where the 'transaction_id' is not null and the 'amount' is positive. Which technique best ensures data quality at this stage?
Apply a post-load DELETE statement to the Silver table after every micro-batch.
Define a separate table for rejected records and use manual SQL queries to move them later.
Use a filter transformation in the Spark Structured Streaming query before writing to the target table.
Filtering during the stream transformation ensures that invalid data is dropped before the commit happens. This is the most efficient method because it eliminates invalid records without needing extra write operations. It keeps the Silver table clean from the beginning, adhering to best practices for production ETL pipelines.
Set the table property 'delta.constraints.check' to filter incoming null values.
A data engineer is migrating legacy batch jobs to Delta Live Tables (DLT). They want to optimize performance for a complex join operation between two large tables. Which TWO strategies should they implement to improve the join efficiency?
Use Z-Ordering on the columns frequently used in the JOIN clause.
Z-Ordering maps multi-dimensional data to one dimension while preserving locality. By Z-Ordering on join keys, Delta Lake clusters related data together in the same files. This significantly reduces the volume of data Spark needs to read during a join, drastically improving query performance for large datasets in production.
Increase the number of partitions to the maximum allowed by the cluster size.
Enable Auto-Optimize for the tables involved in the join.
Auto-Optimize automatically compacts small files during writes, creating larger, more efficient files for readers. This optimization is crucial for join performance, as smaller files force the query engine to perform more I/O operations and metadata lookups, which slow down the shuffle and join phases of the query execution.
Convert the tables to temporary views before executing the join.
Disable the Delta Lake cache to ensure data is read from cloud storage.
You are tasked with handling late-arriving data in a streaming pipeline that performs windowed aggregations. Which approach ensures that the output remains accurate while balancing memory usage?
Increase the state TTL to infinity to ensure no data is ever dropped.
Implement a watermark on the timestamp column and aggregate over a window.
Watermarks allow Spark to understand how much data is allowed to be late. By associating a watermark with a windowed aggregation, Spark can safely discard old state for time windows that have passed the threshold, maintaining both correctness and memory efficiency as the stream processes data over time.
Use the 'complete' output mode to ensure all historical data is reprocessed.
Filter out records with timestamps older than the current processing time.
When implementing a Medallion Architecture, what is the primary purpose of the 'Silver' layer?
To store raw, unmodified data in its original format from source systems.
To provide a cleaned and integrated source of truth for business reporting.
The Silver layer is the trusted source of record. It consolidates data from multiple Bronze sources, enforces quality constraints, and provides a structured schema. This level of refinement ensures that analysts and data science models are based on reliable information, minimizing the time spent on data preparation for downstream tasks.
To act as a sandbox environment for ad-hoc exploration and training.
To aggregate data into pre-calculated metrics for high-performance dashboards.
A team is preparing to optimize their Databricks data transformation pipeline. Which THREE of the following actions are considered best practices for optimizing Delta Lake performance?
Run the OPTIMIZE command to compact small files into larger files.
Compacting small files is critical for Delta Lake performance. Many small files lead to inefficient I/O and excessive metadata operations. The OPTIMIZE command coalesces these into larger, optimally-sized files, which significantly improves read performance for downstream analytical queries, especially when accessing large volumes of historical data.
Partition the data by high-cardinality columns like 'user_id' or 'transaction_id'.
Use Z-Ordering on frequently filtered columns to improve data skipping.
Z-Ordering co-locates data with similar values in the same set of files. When a query filters on these columns, the Delta engine can skip large chunks of data that don't match the criteria. This significantly reduces the amount of data read, resulting in much faster query performance.
Perform a 'VACUUM' operation with a retention period of zero to save costs.
Enable Auto-Compact on the Delta table to automatically manage file sizes during writes.
Auto-Compact reduces the need for manual OPTIMIZE runs by automatically compacting small files during individual write operations. This ensures that the table stays performant without requiring intervention, allowing for a more hands-off and efficient maintenance cycle for streaming or frequently updated Delta tables.
Want more Data Transformation and Modeling practice?
Practice this domainA data engineer needs to configure a Databricks Job to orchestrate a data pipeline that includes a Python task, a SQL task, and a notebook task. The pipeline requires passing a dynamic run identifier from the Python task to the subsequent SQL and notebook tasks. Which mechanism should the data engineer use to achieve this task-to-task dependency parameter passing?
Write the dynamic identifier to a shared Unity Catalog volume and read it inside the downstream tasks.
Use the dbutils.jobs.taskValues.set method in the Python task and reference it downstream using the task value syntax.
The dbutils.jobs.taskValues API allows tasks within a Databricks workflow to securely pass small payloads like run identifiers to downstream tasks. Downstream tasks retrieve these values via task value interpolation, enabling seamless dynamic parameter passing and robust pipeline orchestration.
Store the identifier as an environment variable in the cluster configuration shared by all three tasks.
Modify the global job parameters dynamically via the Databricks REST API from within the running Python script.
A data engineer needs to store structured data in a cloud object storage location while maintaining full ACID guarantees. Which storage format is the foundation of the Databricks Lakehouse architecture that enables this functionality?
JSON
Delta Lake
Delta Lake provides the necessary ACID transactional capabilities, scalable metadata handling, and time travel features required for a robust Lakehouse. It sits on top of existing object storage, ensuring data integrity across complex pipelines while allowing high-performance concurrent access for both streaming and batch data processing workloads.
CSV
Avro
A data engineer is designing a pipeline and needs to ensure that data remains consistent during concurrent read and write operations. Which Databricks feature provides the mechanism to track and validate these operations?
Unity Catalog
Delta Lake Transaction Log
The transaction log is an ordered record of all commits to a Delta table. It acts as the source of truth for the table state, allowing engines to perform optimistic concurrency control. This enables consistent reads and writes, ensuring that data integrity is maintained even during complex multi-writer scenarios.
Databricks SQL
Cluster Manager
Which TWO of the following statements accurately describe the functionality of Unity Catalog within the Databricks Intelligence Platform?
It provides centralized access control, auditing, and lineage for all data and AI assets.
Unity Catalog centralizes permissions, audit logs, and data lineage in one location. This simplifies compliance and security management across multiple workspaces, allowing administrators to apply consistent policies to tables, volumes, and models regardless of which specific Databricks workspace or cloud region an engineer is currently using.
It requires a separate Hive Metastore for each individual Databricks workspace.
It enables seamless cross-workspace data sharing and discovery of assets.
By providing a unified namespace, Unity Catalog allows users to discover and query data assets across different workspaces without needing to replicate data or manage disparate metastores. This promotes collaboration and data democratization while maintaining a high level of security through centralized identity management and object-level permission enforcement.
It is an execution engine used to run heavy machine learning training jobs.
It only supports files stored in Amazon S3 buckets.
When considering the Databricks Intelligence Platform, what is the primary role of the 'Lakehouse' architecture?
To replace cloud object storage with a proprietary proprietary binary format.
To unify the best features of data warehouses and data lakes.
The Lakehouse unifies data warehousing and data lake architectures. It offers the performance and transactional consistency of a warehouse for BI workloads while providing the scalability and support for unstructured data, machine learning, and streaming workloads that are typical of data lakes, all within a single unified platform.
To isolate streaming and batch data processing into distinct storage layers.
To mandate the use of SQL as the only supported programming language.
Refer to the exhibit. A Databricks administrator is using the Unity Catalog JSON policy to manage access. If the 'data_scientist_1' user attempts to execute an 'UPDATE' command on the 'orders' table, what will be the result?
The command will succeed because the user is part of the 'dev_analytics' schema.
The command will fail due to insufficient privileges on the table.
The provided JSON policy only grants 'SELECT' and 'DESCRIBE' permissions. Because Unity Catalog follows a default-deny security posture, the user lacks the required 'UPDATE' or 'MODIFY' permission to alter the contents of the table, causing the Databricks engine to throw an authorization error during the request execution.
The command will succeed because 'SELECT' implies 'UPDATE' in Unity Catalog.
The command will succeed if the cluster is running in Single User mode.
Want more Databricks Intelligence Platform practice?
Practice this domainA data engineering team wants to implement Git integration for their Databricks notebooks. Which workflow is considered the best practice for CI/CD in Databricks Repos?
Export notebooks manually to local machines and commit them via Git CLI.
Use the Databricks REST API to push code updates directly to production notebooks.
Develop code in feature branches, merge via pull requests, and use Databricks Repos to sync production.
This workflow leverages standard DevOps practices like branching and pull requests to ensure code quality through peer reviews. By using Databricks Repos to sync, the production environment stays consistent with the verified main branch, minimizing configuration errors and ensuring that only tested, approved code is executed in production workflows.
Develop code in the production workspace and use Git only for final snapshots.
A team is designing a CI/CD pipeline using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for managing Databricks projects?
DABs automatically provision cloud-native IAM roles for all users.
DABs enable consistent deployment of jobs and pipelines across different environments.
DABs use a declarative YAML structure that allows developers to define jobs, pipelines, and workspace objects once and deploy them consistently across multiple workspaces. This ensures that development, staging, and production environments remain synchronized, preventing the 'works on my machine' syndrome that often plagues manual configuration efforts.
DABs provide built-in visual drag-and-drop tools for building ETL pipelines.
DABs simplify version control for infrastructure and job configurations.
By representing jobs, clusters, and Delta Live Tables pipelines as YAML files, DABs allow these configurations to be stored directly in Git. This enables teams to track changes, review infrastructure modifications through pull requests, and maintain a clear audit trail of who changed which configuration and why.
DABs bypass the need for any authentication tokens during deployment.
Which of the following is a fundamental principle of implementing effective CI/CD for Databricks workflows?
Deploying code from a local laptop directly to the production cluster.
Using environment variables to manage service endpoints and authentication secrets.
Environment variables decouple code from the underlying infrastructure, allowing developers to manage configuration dynamically. This is a best practice in CI/CD as it enables the same codebase to run in any environment by simply injecting the appropriate parameters, reducing the risk of hardcoding secrets or environment-specific values.
Merging all development code into the main branch every few months.
Ignoring unit testing until the code is fully deployed in production.
Why is it important to use Service Principals instead of Personal Access Tokens (PATs) for CI/CD automation in Databricks?
Service Principals are faster to execute than personal tokens.
Service Principals do not require any permission configuration.
Service Principals are not tied to an individual's lifecycle.
Service Principals are machine identities that persist regardless of individual employee status. This prevents the common issue where automated pipelines break due to the expiration or revocation of a user's token. They provide a stable, manageable foundation for long-term CI/CD automation that aligns with enterprise security policies.
Service Principals allow anyone in the organization to run pipelines.
A team is implementing a CI/CD process for their Delta Live Tables (DLT) pipelines. Which THREE of the following practices are recommended to ensure reliable deployment?
Keep all DLT configuration values hardcoded within the pipeline notebook.
Use Git branches to manage feature development and production code.
Branching allows developers to isolate their changes and perform testing without affecting the stable production version. Merging through pull requests ensures that all code changes undergo peer review, which is a critical gatekeeping mechanism for maintaining pipeline stability and preventing unauthorized or faulty code deployments.
Implement automated tests that run against the pipeline before production deployment.
Automated testing verifies that the DLT pipeline logic functions correctly before it is promoted to production. This approach catches integration issues early in the pipeline development cycle, reducing the risk of production failures and ensuring that data quality expectations are met consistently every time a new version is deployed.
Manually update the pipeline source code in the production workspace.
Define pipeline infrastructure using declarative files like JSON or YAML.
Declarative files provide a 'source of truth' for the pipeline's infrastructure. By versioning these files, teams can track the evolution of their pipeline's configuration over time. This infrastructure-as-code pattern simplifies the deployment process by enabling automated tools to apply the exact desired state to any target Databricks workspace.
When promoting code from a development workspace to a production workspace, what is the primary risk of using manual notebook exports?
The notebooks will run significantly slower in the production environment.
The process creates an inconsistent state between development and production.
Manual processes lack the rigor of automated CI/CD pipelines. Differences in libraries, environment settings, or notebook versions often occur when files are manually copied. These inconsistencies lead to code that works in development but fails in production, causing difficult-to-debug errors and potential downtime for critical data tasks.
The Databricks workspace will automatically delete the old notebooks.
Manual exports are prohibited by Databricks security policies.
Want more Implementing CI/CD practice?
Practice this domainA data engineering team needs to ingest millions of small JSON files from an S3 bucket into a Delta Lake table. The solution must provide incremental loading, support schema evolution, and automatically scale to handle increasing file volumes without manual tracking of processed files. Which tool is best suited for this requirement?
The standard Apache Spark DataFrame reader using the .load() method.
The COPY INTO SQL command with the mergeSchema option enabled.
Auto Loader using the cloudFiles source in Structured Streaming.
Auto Loader efficiently handles incremental data loading by tracking the ingestion state through checkpoints. It supports schema inference and evolution, allowing the pipeline to adapt to changes automatically. This makes it the ideal choice for ingesting millions of files from cloud object storage with minimal configuration and maintenance.
A Python loop that iterates through filenames and uses INSERT INTO.
A data engineer is evaluating whether to use Auto Loader or the COPY INTO command for a new ingestion pipeline. Which TWO features are unique to Auto Loader compared to COPY INTO?
The ability to perform incremental loads by tracking processed files.
Native support for schema evolution through schemaEvolutionMode.
Auto Loader provides a dedicated schema evolution feature that can automatically add new columns to the target table when they are detected in the source. This is a significant advantage for handling semi-structured data where the structure might change over time without breaking the primary ingestion pipeline.
Support for ingesting data from cloud object storage like S3 or ADLS.
Integration with cloud-native file notification services for discovery.
Auto Loader can be configured to use cloud-native notification services, such as AWS SNS/SQS or Azure Event Grid, to discover new files. This is more efficient than directory listing for folders containing millions of files, as it avoids the high latency and cost associated with frequent recursive listing.
Requirement for a Delta Lake table as the final destination.
Refer to the exhibit. A data engineer is configuring an Auto Loader stream with the provided options. What will happen if a new JSON file arrives containing a field that is not currently in the target table schema?
The stream will fail and require a manual schema update.
The new field will be added to the target table automatically.
The addNewColumns mode enables Auto Loader to update the table schema dynamically. When a new column is detected in the source files, it is added to the table's metadata and the data is successfully ingested. This allows the pipeline to adapt to upstream changes without stopping or losing data.
The new field will be dropped and only existing columns are kept.
The data for the new field will be moved to a _rescued_data column.
When using Auto Loader to ingest data, why is it recommended to provide a 'schemaLocation' rather than manually defining the entire schema in the code?
To improve the performance of the initial file listing process.
To allow Auto Loader to persist and evolve the schema over time.
The schema location acts as a persistent repository for the schema metadata. By saving the schema here, Auto Loader can detect when the source data structure changes and apply those changes to the target table. This persistence is essential for maintaining the integrity of the incremental loading process across restarts.
To bypass the need for a checkpoint location for the stream.
To encrypt the data schema for security and compliance reasons.
A data engineer is designing a pipeline using Structured Streaming to ingest data into Delta Lake. Which THREE benefits are provided by using checkpoints in this scenario?
Enabling the stream to resume from the exact point of failure.
Checkpoints record the offsets of the data that has been successfully processed. If the streaming job fails or is manually stopped, Spark uses these offsets to determine where to restart the processing. This ensures that no data is skipped and that the pipeline maintains its continuity across different runs.
Providing the ability to recover the previous version of the table.
Ensuring exactly-once processing semantics for the ingestion.
By tracking processed offsets and coordinating with the Delta Lake transaction log, checkpoints ensure that each record is processed and committed exactly once. This prevents the creation of duplicate records in the target table, which is essential for maintaining data quality and accuracy in downstream analytical or reporting applications.
Automatically cleaning up old data files in the target directory.
Storing the state of aggregations across streaming batches.
For stateful operations like aggregations or joins, checkpoints store the intermediate state of the computation. This allows Spark to maintain running totals or windowed calculations across multiple micro-batches, ensuring that the results remain accurate even if the stream is interrupted and subsequently restarted at a later time.
While ingesting CSV files using Auto Loader, a data engineer notices that some records have malformed data that does not match the inferred schema. How can the engineer capture these records without failing the entire ingestion stream?
By setting the 'mode' option to 'FAILFAST' in the reader.
By enabling the '_rescued_data' column in the Auto Loader configuration.
When the rescued data column is enabled, Auto Loader automatically places any data that cannot be parsed into the expected schema into a special JSON column. This allows the rest of the record's valid fields to be processed normally while preserving the malformed content for future debugging.
By using a TRY_CAST function in a transformation after the read.
By increasing the 'maxFilesPerTrigger' to handle more errors.
Want more Data Ingestion and Loading practice?
Practice this domainA Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?
Increase the cluster's worker node instance type with more memory.
Scaling up to worker nodes with higher memory capacity provides the Spark executors with more heap space. This allows them to process larger partitions and handle complex shuffle operations without spilling to disk or triggering OOM exceptions. This is the most direct hardware-level fix for memory-intensive join operations.
Decrease the number of partitions in the Spark cluster.
Disable the spark.sql.autoBroadcastJoinThreshold configuration.
Broadcasting an excessively large table triggers OOM errors if the table exceeds the executor memory limit. By setting this threshold to a lower value or disabling it, you force Spark to perform a sort-merge join instead of a broadcast join, which is more robust for large datasets that do not fit in memory.
Enable dynamic allocation for the cluster.
Use the cache() method on every DataFrame in the pipeline.
Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?
The Delta Lake History logs.
The Spark UI SQL tab.
The Spark UI's SQL tab provides detailed information on how Spark parses, optimizes, and executes a query. It allows engineers to inspect the physical plan, identify time-consuming operators, and check for issues such as skewed joins or excessive data shuffling, making it the primary tool for query troubleshooting.
The Cluster Metrics dashboard.
The Data Explorer.
A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?
The trigger interval is set to a time shorter than the processing time of the batch.
If the time taken to process a batch exceeds the trigger interval, the streaming job cannot keep up. This leads to a queue of pending batches, causing the watermark and processing time to drift further behind, effectively indicating that the cluster is undersized for the current workload volume.
Lack of proper watermarking on stateful aggregations.
Without watermarks, the state store grows indefinitely as Spark keeps all keys in memory. This causes increasingly longer garbage collection pauses and slower processing times for each subsequent batch, eventually leading to massive lag as the system struggles to manage the bloated state across nodes.
Using a fixed-size cluster with no auto-scaling enabled.
Inefficient shuffling due to data skew.
Data skew forces specific executors to process significantly more data than others, becoming a bottleneck for the entire stream. Since the micro-batch must wait for all tasks to complete, the slowest task determines the batch latency, causing the job to fall behind regardless of the total cluster resources.
Using too few partitions in the input source.
An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?
Use a Broadcast Hash Join.
A broadcast join sends the smaller table to all worker nodes. This eliminates the need to shuffle the large table, which is the most expensive part of a join operation in Spark. This strategy is highly effective when one side of the join is significantly smaller than the other.
Implement a Shuffle Hash Join.
Increase the number of shuffle partitions.
Enable Z-Ordering on the join key.
Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?
DELETE FROM table_name WHERE date < current_date() - 30;
TRUNCATE TABLE table_name;
VACUUM table_name RETAIN 30 DAYS;
VACUUM is the specific command used to delete data files that are no longer part of the Delta table's current state and are older than the specified retention threshold. This operation is necessary to reclaim storage space in cloud object storage and maintain compliance with data management policies.
DROP TABLE table_name;
Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?
The Spark UI.
The Ganglia UI.
Ganglia is the built-in monitoring tool in Databricks that provides deep, time-series insights into cluster-level health metrics like CPU load, memory utilization, and network traffic. It is the primary resource for troubleshooting hardware-level bottlenecks and determining if a cluster is appropriately sized for the workload it is executing.
The Query Profile.
The Databricks Job Run History.
Want more Troubleshooting, Monitoring, and Optimization practice?
Practice this domainA data engineer needs to configure a Databricks Job containing multiple tasks where downstream tasks should only execute if all upstream parent tasks complete successfully. Which task dependency setting should be configured?
Configure a retry policy with exponential backoff on the parent task to guarantee eventual success.
Set up a global workspace webhook to trigger downstream tasks manually after the parent completes.
Add the upstream tasks as parents of the downstream task directly within the job configuration graph.
Declaring upstream tasks as parents of the downstream task in the job graph creates explicit dependencies, so Databricks triggers the downstream task only after every parent completes successfully. This directly enforces the all-upstream-success condition the scenario requires.
Enable concurrent run suppression on the job level to prevent overlapping workflow executions.
Which TWO of the following statements are true regarding the behavior and capabilities of Databricks Jobs parameters and values?
Task values can be used to pass small amounts of data, such as counts or status codes, from one task to a downstream task.
The dbutils.jobs.taskValues API allows tasks to set key-value pairs that downstream tasks can retrieve. This facilitates dynamic branching and parameter passing based on runtime metrics generated during the execution of upstream workflow steps.
Job parameters defined at the workflow level can be overridden when triggering a run via the Databricks CLI or REST API.
When triggering job executions using automation tools like the CLI or REST API, you can supply custom parameters that override default values configured in the job definition. This provides flexibility for CI/CD pipelines and external orchestrators.
Job parameters are automatically persisted in a Unity Catalog volume after the workflow completes successfully.
Notebook tasks cannot accept parameter values passed from the parent job orchestration configuration.
Task values support passing large dataframes containing millions of rows between dependent tasks.
A data engineer has configured a Databricks Job with multiple dependent tasks forming a linear pipeline. Task A extracts data, Task B transforms it, and Task C loads it into a gold table. The pipeline runs daily. The team notices that if Task B fails due to an intermittent schema validation issue, the entire job run fails, but they want Task C to execute conditionally only if Task B succeeds, while alerting the on-call engineer immediately upon any failure. How should the task dependencies and conditional execution be configured?
Set Task C to run on failure of Task B to capture the exception state before alerting.
Convert Task B and Task C into a single notebook task and handle exceptions internally with try-except blocks.
Ensure Task C has a dependency pointing exclusively to Task B, and configure a job-level email notification for failures.
Establishing a direct parent-child dependency between Task B and Task C guarantees that the load step never runs prematurely. Adding job-level failure alerts ensures immediate notification to the engineering team without compromising pipeline DAG semantics.
Disable the task dependency between Task B and Task C and run them in parallel with different cluster configurations.
A data engineer needs to configure a Databricks Job containing three distinct tasks: ingest, transform, and report. The transform task must only execute if the ingest task completes successfully, but the report task should execute regardless of whether the transform task succeeds or fails. How should the task dependencies be configured?
Configure transform to depend on ingest, and configure report to depend on transform.
Configure ingest to depend on transform, and configure report to depend on ingest.
Configure transform to depend on ingest, and do not include transform as a dependency for report.
Setting transform to depend on ingest ensures the data is loaded before transformation begins. Omitting transform from report's dependencies allows the reporting workflow to run independently of the transformation task's success or failure state.
Configure both transform and report to depend on ingest simultaneously.
A data engineer wants to ensure that a Databricks Job task only runs if the preceding task completes successfully, but needs to add a specific timeout threshold for this individual task. Where should this configuration be applied?
At the job level settings.
Within the individual task definition.
Task definitions contain specific configurations like timeout_seconds, retries, and depends_on arrays. Setting the timeout here ensures that only this specific task is terminated if it exceeds the threshold, while allowing downstream tasks to potentially trigger or fail gracefully based on the defined job dependency graph.
Inside the cluster configuration JSON.
Using a global Workspace policy.
A data engineer is designing a Databricks Job workflow. Which TWO of the following are valid ways to trigger a Databricks Job?
Using the 'Run now' button in the UI.
The 'Run now' feature is the primary way to manually execute a job immediately for testing or ad-hoc data processing requirements. It bypasses the schedule and triggers an instance of the job using the existing configuration, making it indispensable for troubleshooting or re-running failed tasks.
Directly editing the Python source code file.
Using the Databricks Jobs REST API.
The Jobs API allows for programmatic triggers, enabling integration with external orchestrators like Airflow or CI/CD pipelines. By sending a POST request to the /api/2.1/jobs/run-now endpoint, engineers can reliably initiate workflows based on external events, ensuring data freshness and consistency across distributed enterprise systems.
Changing the cluster node type.
Modifying the job owner permissions.
Want more Working with Lakeflow Jobs practice?
Practice this domainA data engineer needs to share a Delta table managed by Unity Catalog with external partners who do not have access to the Databricks workspace. Which feature should be used to securely grant read-only access to this table without replicating the data?
Configure cross-account IAM role assumption to grant direct AWS S3 bucket read permissions to the external partners.
Set up a Databricks SQL endpoint and create individual workspace user accounts for every external partner.
Create a Delta Sharing share, add the table to the share, and grant access to a recipient object configured for the partners.
Delta Sharing is an open protocol that lets recipients read shared tables through a recipient object without workspace access or data replication. This satisfies the constraint of secure read-only external sharing while keeping a single governed copy.
Export the Delta table to CSV format and upload the files to an external SFTP server on a nightly schedule.
A data engineer is configuring Unity Catalog governance for a multi-department organization. Which TWO actions require the metastore admin or catalog owner to have a workspace-independent metastore assigned? (Choose TWO)
Creating a new catalog inside the Unity Catalog metastore.
Binding a workspace to a Unity Catalog metastore to enable catalog-level data access.
Workspace binding enforces security isolation by mapping specific workspaces to a Unity Catalog metastore. This administrative configuration ensures that data assets governed by the metastore are only accessible through authorized workspaces, requiring proper metastore administration.
Granting the SELECT privilege on a specific schema to a group of data analysts.
Configuring the root storage location for the metastore using a secure cloud storage container.
Defining the root storage location establishes the foundational storage layer for managed tables within the metastore. This critical setup requires high-level administrative permissions and infrastructure provisioning before any workspace users can interact with managed data.
Defining a custom dynamic view with column-level masking functions for auditing.
An organization requires that all data access logs across multiple Databricks workspaces be captured and sent to a centralized security information and event management (SIEM) system. Which Databricks feature should be configured to capture these audit events?
Enable Spark event logs in the cluster configuration UI for every running compute cluster.
Configure diagnostic log delivery to stream account and workspace audit logs to cloud storage or event hubs.
Diagnostic log delivery streams account and workspace audit logs to cloud storage or event hubs, giving the centralised feed the SIEM requires. This directly satisfies the requirement to capture audit events across multiple workspaces into one security information and event management system.
Run a nightly notebook that queries the system.information schema for active user sessions.
Install a Python logging package on every cluster driver node to intercept SQL queries.
A data engineer is migrating legacy tables to Unity Catalog. Which TWO of the following statements regarding the transition to three-level namespace (catalog.schema.table) are true?
Three-level namespace applies to all Unity Catalog objects.
Unity Catalog enforces a strict hierarchy consisting of catalog, schema, and table or view. This structure provides a unified way to manage permissions and lineage across different workspaces. Every managed or external object must adhere to this naming convention to be correctly registered within the Unity Catalog metadata store.
Hive Metastore objects are automatically migrated to Unity Catalog.
Permissions at the catalog level cascade to all child schemas and tables.
Unity Catalog uses an inheritance model where privileges granted at a higher level, such as the catalog, automatically apply to all schemas and tables contained within it. This design significantly reduces the administrative burden of managing individual permissions for thousands of database objects across an entire enterprise organization.
External locations are not required for three-level namespace tables.
Three-level namespace is optional in Unity Catalog.
Refer to the exhibit. A data engineer attempts to run the GRANT statement above in a Databricks SQL query editor. What is the most likely reason for failure?
The syntax of the GRANT statement is incorrect.
The table does not exist.
The table must be referenced using a three-level namespace.
Unity Catalog requires the fully qualified name (catalog.schema.table). Providing only the table name causes the operation to fail because the metastore cannot resolve the resource context. This forces explicit identification of the data, which is a security best practice to prevent unauthorized or unintended access to tables.
The group 'data_scientists' does not exist.
Which feature in Unity Catalog is primarily used to track data movement and transformation history for compliance and auditing?
Audit Logs
Data Lineage
Data lineage captures the dependencies between tables, views, and notebooks. It shows the flow of data through transformations, which is essential for auditability. By tracking these relationships, data engineers and security officers can verify the provenance of data and fulfill regulatory requirements regarding data processing and handling.
Delta Sharing
Catalog Explorer
Want more Governance and Security practice?
Practice this domainThe Databricks-DE-Assoc exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 7 domains: Data Transformation and Modeling, Databricks Intelligence Platform, Implementing CI/CD, Data Ingestion and Loading, Troubleshooting, Monitoring, and Optimization, Working with Lakeflow Jobs, Governance and Security. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Databricks Databricks-DE-Assoc exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.