Databricks-DE-Pro Data Ingestion and Acquisition Practice Question
Your organization is ingesting sensitive PII data. You need to ensure that personal identifiers are masked during the ingestion process before they are stored in the Bronze layer of your Medallion architecture. What is the best practice for this?
⚠ Common exam trap
Students mistakenly think PII data should be masked downstream in the Gold layer for final reporting, overlooking the critical compliance requirement to secure raw data early.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply masking logic within the initial streaming ingestion transformation.
Implementing masking at the ingestion layer using Delta Live Tables (DLT) expectations or standard Spark transformations ensures that sensitive data is never written in plaintext to the Bronze layer. Applying transformations during the 'Acquisition' phase is a critical security practice, ensuring that governance requirements are met before data becomes available to analysts, thereby reducing the risk of accidental exposure and maintaining data privacy compliance across the pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Mask the data using Delta Lake column masking after the data reaches the Silver layer.
Why it's wrong here
Masking at the Silver layer leaves the sensitive data exposed in the Bronze layer. This violates the principle of least privilege, as anyone with access to the raw ingestion table would be able to view the unmasked PII. Security should be applied as early as possible in the pipeline.
- ✓
Apply masking logic within the initial streaming ingestion transformation.
Why this is correct
Applying masking logic during the initial ingestion transformation (e.g., using withColumn and sha2 or similar functions) ensures that PII is protected before it is ever committed to storage. This maintains data privacy from the moment the data enters the ecosystem, preventing sensitive information from ever reaching the persistent Bronze table.
- ✗
Configure Unity Catalog to mask columns only for specific users.
Why it's wrong here
Unity Catalog dynamic masking is an effective governance tool for query-time control. However, it does not remove the actual data from the underlying storage. If the data is ingested in plaintext, it remains vulnerable to anyone with direct read access to the underlying storage files on your cloud provider.
- ✗
Use a post-ingestion job to delete PII columns.
Why it's wrong here
Running a post-ingestion job leaves a window of time where sensitive data is stored in the Bronze layer in its raw, unmasked form. This is a security risk and is inefficient compared to applying transformations inline during the ingestion process, which avoids writing the sensitive data to disk entirely.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.