Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

A Data Engineer is building a Delta Live Tables (DLT) pipeline to ingest raw JSON data. They need to ensure that records missing the required 'user_id' field are dropped while simultaneously capturing these discarded records in a separate table for auditing purposes. Which approach achieves this in DLT?

⚠ Common exam trap

Candidates often search for a single DLT command that drops and saves records simultaneously. They fail to realize that DLT requires two distinct logic paths to separate valid and invalid data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define two separate tables in the pipeline, one filtering for valid records and one for invalid records using the NOT condition.

To handle data quality in DLT, the 'expect_violation_or_drop' constraint is not a standard clause. Instead, the 'expect_or_drop' constraint removes invalid records, but does not preserve them. The correct architectural pattern involves using a separate pipeline or query that filters for the inverse condition (where 'user_id' is null) and writes those records to an 'expect_all_fail' target table, maintaining strict lineage and auditing for schema non-compliance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply the 'expect_or_drop' constraint and use a trigger to log dropped records.

    Why it's wrong here

    Delta Live Tables does not provide a native trigger mechanism for logging dropped records during the execution of an 'expect_or_drop' constraint. This approach would result in silent data loss for the records being filtered out, failing the audit requirement defined in the business scenario for this specific pipeline.

  • ✗

    Use the 'expect_or_fail' constraint to halt the pipeline and flag the error.

    Why it's wrong here

    The 'expect_or_fail' constraint stops the entire pipeline execution when a violation is detected. This contradicts the requirement to continue processing valid records while merely separating the invalid ones for auditing. It creates a bottleneck and prevents the successful ingestion of valid data points during batch execution.

  • ✓

    Define two separate tables in the pipeline, one filtering for valid records and one for invalid records using the NOT condition.

    Why this is correct

    By defining two tables, you leverage the declarative nature of DLT to materialize valid and invalid datasets concurrently. Using the inverse logical condition for the audit table ensures that all records are accounted for, meeting both the ingestion requirement and the audit policy without halting the pipeline's progress.

  • ✗

    Configure a DLT pipeline to use 'expect_all_drop' to automatically split the data stream.

    Why it's wrong here

    There is no 'expect_all_drop' constraint in Delta Live Tables syntax. The available constraints are 'expect', 'expect_or_drop', and 'expect_or_fail'. Attempting to use non-existent syntax will cause the pipeline initialization to fail, preventing the deployment of the data transformation logic entirely within the Databricks environment.

About these practice questions

Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.