Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

You are migrating a legacy CSV-based ETL process to Databricks. The source CSV files contain inconsistent date formats. Which approach provides the most scalable way to handle these inconsistencies during the bronze-to-silver transformation?

⚠ Common exam trap

Candidates often attempt to fix inconsistent date formats using single-format parsers or raw string manipulation, ignoring functions that accept multiple fallback format patterns.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the 'to_timestamp' function with an array of acceptable formats to parse the column.

Using Spark's `to_timestamp` with multiple format strings or a custom UDF is the standard, scalable way to handle format drift. By applying this logic in the Silver layer, you preserve the raw data in the Bronze layer while ensuring the refined data is standardized. This strategy follows the Medallion architecture pattern, allowing for lineage tracking and the ability to reprocess data if requirements change or better parsing logic is developed.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert all columns to string type to avoid parsing errors, and handle formatting in the visualization tool.

    Why it's wrong here

    Storing dates as strings is a bad practice that limits the functionality of time-series analysis, partitioning, and date arithmetic. It shifts the burden of cleaning to downstream users, increasing the likelihood of errors and reducing performance, as tools cannot efficiently filter or group data without proper date types.

  • ✗

    Use the 'from_unixtime' function with a hardcoded string format for all records.

    Why it's wrong here

    Hardcoding a single format string will fail for any records that deviate from that specific format, causing the pipeline to crash or generate nulls. This approach is brittle and cannot handle the 'inconsistent' date formats described in the scenario, making it unsuitable for real-world messy data scenarios.

  • ✓

    Use the 'to_timestamp' function with an array of acceptable formats to parse the column.

    Why this is correct

    The 'to_timestamp' function in Spark SQL accepts a format string or an array of formats. This allows the engine to attempt parsing against multiple patterns, which is the most robust and performant way to handle format inconsistencies without writing complex, slow-running row-based logic or custom UDFs.

  • ✗

    Delete all rows with invalid date formats using a 'drop' expectation in a DLT pipeline.

    Why it's wrong here

    Automatically dropping data loses potentially valuable information. While it cleans the dataset, it does not solve the underlying problem of data ingestion. A robust pipeline should attempt to parse and standardize the data first, and only quarantine or drop if the data is genuinely unrecoverable or corrupted.

About these practice questions

One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.