Courseiva

Databricks-DE-Pro · topic practice

Data Transformation, Cleansing, Quality practice questions

This domain covers building reliable transformations on Databricks: cleansing and validating data, handling skew and nested types, and enforcing quality in Delta Lake and Lakeflow Spark Declarative Pipelines. Questions present concrete scenarios—self-joins on skewed tables, clickstream STRUCT columns, JSON sensor ingestion, and MERGE constraint violations—and ask you to choose the correct technique or predict default behavior.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Data Transformation, Cleansing, Quality

What the exam tests

What to know about Data Transformation, Cleansing, Quality

A candidate must design transformations that cleanse, validate, and join data correctly on Databricks. The most important thing is knowing how Delta Lake and Lakeflow pipelines handle invalid records by default, and applying skew mitigation like salting when joins on large tables are involved.

Applying salting and skew hints to optimize self-joins on large skewed Delta tables

Flattening and casting nested STRUCT fields from Delta clickstream payloads in Databricks SQL

Using Lakeflow Spark Declarative Pipelines expectations to drop or quarantine invalid records

Predicting Delta Lake default behavior when MERGE rows violate CHECK or NOT NULL constraints

Watch out for

Common Data Transformation, Cleansing, Quality exam traps

  • ▸Assuming Delta MERGE silently drops bad rows; by default constraint violations raise errors and can abort the transaction unless handled.
  • ▸Repartitioning or broadcasting a skewed self-join without salting the key, which leaves straggler tasks and does not fix skew.
  • ▸Treating nested STRUCT fields as flat columns instead of using dot notation or explode, causing schema and analysis errors.

Practice set

Data Transformation, Cleansing, Quality questions

20 questions · select your answer, then reveal the explanation

A Data Engineer is building a pipeline using Delta Live Tables (DLT) to clean IoT sensor data. Which TWO of the following statements regarding the implementation of Expectations are correct?

Refer to the exhibit. A pipeline job is failing because of a schema mismatch between the source data and the Delta table. Which solution allows the pipeline to succeed without manually changing the source data or the existing table schema?

Exhibit

Error: [DELTA_SCHEMA_MISMATCH] The schema of the incoming data does not match the table schema. Expected: [id: LONG, name: STRING, ts: TIMESTAMP]. Found: [id: LONG, name: STRING, ts: STRING].

You are optimizing a pipeline that processes high-volume JSON data. Which THREE techniques will improve the performance and quality of the transformation layer?

A data engineering team is optimizing a massive Delta Lake table partitioned by date and clustered by customer_id using Liquid Clustering. The table experiences frequent updates and deletes based on streaming CDC feeds. Which underlying Delta Lake feature allows this pattern to perform efficiently without causing file compaction bottlenecks?

Refer to the exhibit. A Data Engineer is reviewing a DLT configuration file. What is the impact of the 'on_violation' setting on the pipeline?

Exhibit

{
  "constraint": "valid_email",
  "expect": "email IS NOT NULL AND email LIKE '%@%'",
  "on_violation": "fail_update"
}

A Data Engineer needs to perform a complex data cleaning operation that involves joining a streaming table with a large, static lookup table. The lookup table is updated daily via a separate batch process. What is the most efficient way to ensure the join is performant and consistent within a DLT pipeline?

A data engineer is building a Delta Live Tables pipeline that ingests raw JSON events and must produce a cleansed Silver table. The pipeline must handle malformed records without failing, capture metrics about data quality, and allow downstream analysts to query only valid records. Which two actions should the engineer take? (Choose two.)

A Data Engineer is working with a Delta table that contains a column `email` with inconsistent casing and leading/trailing spaces. The engineer needs to standardize the email addresses by trimming spaces and converting to lowercase. Which SQL expression correctly transforms the `email` column in a SELECT statement?

A data engineer is implementing a data quality check on a streaming Delta table that receives late-arriving events. The engineer needs to detect records where the event timestamp is more than 2 hours behind the current watermark. Which approach correctly identifies such records for quarantine?

A Data Engineer is building a Databricks job to cleanse a large dataset. The engineer needs to identify and handle duplicate records, missing values, and outliers. Which TWO techniques are most appropriate for addressing these data quality issues in a scalable way? (Choose two.)

A Data Engineer is tasked with cleansing a large Delta table containing customer records. The table has columns `customer_id`, `email`, `signup_date`, and `country`. The engineer needs to identify and correct data quality issues such as invalid email formats, future signup dates, and inconsistent country codes. Which TWO approaches are most appropriate for performing these cleansing operations in Databricks? (Choose two.)

A Data Engineer is implementing a data quality framework in Databricks using Delta Live Tables (DLT). The framework must validate incoming data against multiple rules, track metrics, and ensure that critical violations halt the pipeline. Which TWO actions should the engineer take to meet these requirements? (Choose two.)

A data engineer is using Delta Live Tables to build a pipeline that ingests CSV files with a known schema. The source files occasionally contain extra columns that are not in the schema. The engineer wants the pipeline to ignore these extra columns without failing. Which option should be set in the DLT table definition?

A data engineer is building a Delta Live Tables (DLT) pipeline to ingest and cleanse streaming data from a Kafka topic. The pipeline must ensure data quality by dropping records that fail validation rules and recording metrics about dropped records. Which two features should the engineer use to achieve this? (Choose two.)

A Data Engineer is implementing a data quality framework for a Delta table that receives streaming updates. The engineer needs to ensure that any record with a negative `amount` is not written to the table, but also wants to capture these invalid records in a separate quarantine table for later analysis. The pipeline must not fail on invalid records. Which approach using Delta Live Tables achieves this?

A Data Engineer is designing a Delta Live Tables pipeline that ingests data from a Kafka stream. The pipeline must ensure that records with a null 'transaction_id' are dropped, and records where 'amount' is less than 0 are flagged but still processed. Additionally, the pipeline should quarantine records that fail both conditions. Which combination of DLT expectations and pipeline settings achieves this?

A Data Engineer is using Databricks SQL to cleanse a table `sales` that has a column `amount` stored as a string, but it contains numeric values with dollar signs and commas. The engineer needs to convert this column to a decimal type for analysis. Which expression correctly casts the cleaned string to DECIMAL(10,2)?

A data engineer is tasked with cleaning a dataset in Databricks. The dataset contains a column 'email' with some null values. The engineer wants to replace nulls with a default value 'unknown@example.com'. Which PySpark function should be used?

A Data Engineer needs to enforce a NOT NULL constraint on a specific column in a Delta table while maintaining the ability to perform high-performance streaming writes. Which approach is the most efficient and native method to ensure this data quality requirement?

Refer to the exhibit. A Data Engineer is attempting to merge data into a table with these constraints defined. If the incoming batch contains rows that violate these rules, what is the default behavior of the Delta Lake engine during the merge operation?

Exhibit

{
  "type": "Table",
  "constraints": {
    "check": "price > 0",
    "not_null": ["id", "transaction_date"]
  }
}

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Data Transformation, Cleansing, Quality sessions

Start a Data Transformation, Cleansing, Quality only practice session

Every question in these sessions is drawn from the Data Transformation, Cleansing, Quality domain — nothing else.

Related practice questions

Related Databricks-DE-Pro topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Pro exam test about Data Transformation, Cleansing, Quality?
A candidate must design transformations that cleanse, validate, and join data correctly on Databricks. The most important thing is knowing how Delta Lake and Lakeflow pipelines handle invalid records by default, and applying skew mitigation like salting when joins on large tables are involved.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Data Transformation, Cleansing, Quality questions in a focused session?
Yes — the session launcher on this page draws every question from the Data Transformation, Cleansing, Quality domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Pro topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Pro question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Pro exam covers. They are not copied from any real exam or dump site.