Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

When designing a Data Quality framework in Databricks, what is the recommended approach for handling 'quarantined' records?

⚠ Common exam trap

Candidates frequently suggest deleting invalid records outright. They often overlook the requirement for auditing and reprocessing, which necessitates moving data to a quarantine table rather than simply discarding it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Move invalid records to a 'quarantine' table with an additional column indicating the error.

Storing invalid records in a separate, dedicated table allows for auditing and correction without blocking the main pipeline. This ensures high throughput for valid data while providing a clear path for reprocessing bad data once the underlying issue is fixed. It is a fundamental pattern for resilient data engineering, ensuring that data is never silently dropped and that stakeholders maintain confidence in the data's reliability.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Delete the bad records and notify the team via an automated Slack notification.

    Why it's wrong here

    Simply deleting records is destructive and hides data quality issues. While notification is helpful, it doesn't solve the problem of how to recover the lost data. Quarantining is preferred because it preserves the original records for later analysis and correction, which is critical for audit and business continuity.

  • ✓

    Move invalid records to a 'quarantine' table with an additional column indicating the error.

    Why this is correct

    This approach is the gold standard for robust ETL. By tagging records with the reason for failure, engineers can easily analyze the patterns of error and fix the source issues. It maintains a clean primary table while simultaneously building a valuable dataset of quality issues for analysis.

  • ✗

    Update the records in place by setting the invalid fields to NULL.

    Why it's wrong here

    Updating records to NULL destroys information and makes it difficult to understand why the data was invalid in the first place. It also makes it impossible to reprocess those records later because the original values are permanently lost during the 'clean' operation, which is not an acceptable data practice.

  • ✗

    Stop the entire pipeline execution to ensure no bad data reaches the final target.

    Why it's wrong here

    Stopping the pipeline is a 'nuclear' option that causes unnecessary downtime and data delays. In a modern data environment, it is better to isolate bad data and continue processing the good data, ensuring that critical analytical dashboards remain updated while the specific quality issues are addressed in parallel.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.