Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question
A Data Engineer is building a Lakeflow Spark Declarative Pipelines pipeline that ingests JSON sensor events. The pipeline must drop records where the `sensor_id` is NULL, ensure that `event_time` is not in the future, and continue processing without failing the update. Which combination of expectations should be used?
⚠ Common exam trap
A common mix-up: candidates confuse `expect_all` (which only logs metrics) with `expect_or_drop` (which actually removes records), or assuming that `expect_or_fail` is needed to enforce data quality.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use `@dp.expect_or_drop("valid_sensor", "sensor_id IS NOT NULL")` and `@dp.expect_or_drop("valid_time", "event_time <= current_timestamp()")` on the dataset.
The pipeline must drop records that violate the conditions and continue processing. `expect_or_drop` is designed for this: it discards invalid rows and allows the update to succeed. Using `expect_or_fail` would halt the pipeline, and `expect_all` would not remove the bad records, leaving them in the target table.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use `@dp.expect_all({"valid_sensor": "sensor_id IS NOT NULL", "valid_time": "event_time <= current_timestamp()"})` on the dataset.
Why it's wrong here
`expect_all` only records metrics for violations but does not drop the offending records. The bad rows would still be written to the target table, violating the requirement to drop them. This decorator is for monitoring only, not for enforcing data quality by removing invalid data.
- ✗
Use `@dp.expect_or_fail("valid_sensor", "sensor_id IS NOT NULL")` and `@dp.expect_or_drop("valid_time", "event_time <= current_timestamp()")` on the dataset.
Why it's wrong here
Mixing `expect_or_fail` for the sensor_id check means that if any record has a NULL sensor_id, the entire pipeline update fails. The requirement is to drop invalid records and continue, so using `expect_or_fail` is inappropriate for the first condition. Only `expect_or_drop` for both would satisfy the scenario.
- ✗
Use `@dp.expect_all_or_fail({"valid_sensor": "sensor_id IS NOT NULL", "valid_time": "event_time <= current_timestamp()"})` on the dataset.
Why it's wrong here
This decorator will fail the entire pipeline update if any record violates either expectation. The requirement is to continue processing, not to stop. While it enforces data quality, it does not drop the bad records and allow the pipeline to proceed, so it fails the scenario's need for non-blocking handling.
- ✓
Use `@dp.expect_or_drop("valid_sensor", "sensor_id IS NOT NULL")` and `@dp.expect_or_drop("valid_time", "event_time <= current_timestamp()")` on the dataset.
Why this is correct
These expectations drop records that violate the conditions while allowing the pipeline update to complete successfully. `expect_or_drop` is the correct decorator when you want to discard bad records and not fail the pipeline. Both conditions are expressed as SQL expressions, which are evaluated per row. The pipeline continues processing, and dropped records are tracked in event logs and metrics.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.