Databricks-DE-Pro Data Transformation, Cleansing, Quality Practice Question
You are tasked with handling PII (Personally Identifiable Information) in your data pipeline. Which approach is best for protecting this data while maintaining the ability to perform analytics?
⚠ Common exam trap
Candidates frequently assume that dropping PII columns is the only compliance method, forgetting that cryptographic hashing allows analytics while protecting sensitive data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Hash the PII columns using a salted, non-reversible cryptographic hash function.
Hashing or tokenizing PII allows for unique identification without exposing sensitive information. This is a standard practice for compliance (e.g., GDPR, CCPA). By using consistent hashing, you maintain the ability to join tables or perform group-bys on the hashed key while ensuring that unauthorized users cannot reverse-engineer the sensitive original values, balancing security with functional analytical utility.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Simply drop the PII columns during the Bronze-to-Silver transformation.
Why it's wrong here
Dropping PII entirely removes the ability to perform vital analytics, such as customer journey tracking or segment-based analysis. While it is secure, it is often not a viable business solution because it destroys the value of the data for downstream stakeholders who require that level of detail.
- ✗
Encrypt the PII data using a shared key stored in the pipeline code.
Why it's wrong here
Hardcoding keys in pipeline code is a security vulnerability. If the code is accessed, the encryption is compromised. Additionally, encryption/decryption overhead can slow down pipelines, and managing shared keys manually is not scalable or secure compared to using dedicated services like Azure Key Vault or AWS KMS.
- ✓
Hash the PII columns using a salted, non-reversible cryptographic hash function.
Why this is correct
Hashing with a salt provides a secure, non-reversible way to mask PII while maintaining referential integrity. This allows analysts to group data by the hash key without seeing the actual sensitive values. It is a highly recommended practice for balancing data utility with strict regulatory and privacy requirements.
- ✗
Use a public, non-salted MD5 hash function to ensure consistency across teams.
Why it's wrong here
MD5 is cryptographically weak and susceptible to rainbow table attacks. Using a public, unsalted hash makes it easy for attackers to pre-compute hashes and map them back to real-world identifiers, effectively rendering the anonymization process useless and leaving the organization vulnerable to privacy breaches and regulatory non-compliance.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.