Courseiva

Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question

A data engineer is using PySpark to cleanse a large dataset of customer records. The DataFrame `df` contains a string column `phone` with values like '123-456-7890', '(123) 456-7890', and '1234567890'. The engineer needs to standardize these to digits only (e.g., '1234567890'). Which transformation should be used?

⚠ Common exam trap

The trap here is assuming that translate can remove characters by mapping them to an empty string, but translate requires equal-length mapping strings and will not work as intended for removal.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

df.withColumn('phone_clean', regexp_replace('phone', '[^0-9]', ''))

To standardize phone numbers to digits only, all non-digit characters must be removed. The regexp_replace function with the pattern '[^0-9]' efficiently replaces any character that is not a digit with an empty string, resulting in a clean digit-only string. This method is robust and handles various formats without complex string manipulation. Other functions like translate or split do not directly produce the desired output.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    df.withColumn('phone_clean', split('phone', '[^0-9]'))

    Why it's wrong here

    The split function divides a string into an array of substrings based on a delimiter pattern. Using a non-digit pattern would split the phone number into multiple parts, resulting in an array of digits and empty strings, not a single concatenated digit string. This does not achieve the standardization goal and would require further processing to concatenate the array elements.

  • ✗

    df.withColumn('phone_clean', translate('phone', '-() ', ''))

    Why it's wrong here

    The translate function replaces characters one-to-one based on a mapping string. Here, it would map '-' to nothing, '(' to nothing, etc., but the mapping string lengths must match. Providing a shorter replacement string would cause an error or unexpected behavior because translate requires both strings to be of equal length. This approach is error-prone and does not handle all non-digit characters comprehensively.

  • ✓

    df.withColumn('phone_clean', regexp_replace('phone', '[^0-9]', ''))

    Why this is correct

    The regexp_replace function replaces all non-digit characters with an empty string, effectively removing hyphens, parentheses, and spaces. This standardizes the phone numbers to a contiguous digit string, which is the desired outcome. It operates on the column 'phone' and creates a new column 'phone_clean' without modifying the original data, aligning with typical cleansing practices.

  • ✗

    df.withColumn('phone_clean', trim('phone'))

    Why it's wrong here

    The trim function only removes leading and trailing whitespace from a string. It does not remove internal hyphens, parentheses, or spaces, so '123-456-7890' would remain unchanged. This fails to standardize the phone numbers to digits only, as required. Additional transformations would be necessary to achieve the desired cleansing.

About these practice questions

This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.