Databricks-GenAI-Assoc Data Preparation Practice Question
You are preparing a large text corpus for fine-tuning a generative AI model. The corpus is stored in a Delta table with columns: doc_id, raw_text, and metadata. You need to create a cleaned dataset that removes personally identifiable information (PII) and normalizes whitespace, while preserving document boundaries for training. Which two actions should you perform to achieve this in a scalable and maintainable way? (Choose two.)
⚠ Common exam trap
The trap here is thinking that any AI function can clean text, when only ai_mask targets PII and regex handles whitespace; other AI functions like sentiment or classification do not perform the required transformations.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply a regular expression to replace all whitespace sequences with a single space in the raw_text column.
ai_mask redacts PII in place, and regexp_replace normalizes whitespace, together achieving the cleaning goals at scale. Both are declarative, repeatable, and preserve document boundaries. Sentiment filtering, manual inspection, and classification are either irrelevant or destructive to the dataset, and they do not address PII removal and whitespace normalization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the Delta table to Parquet files and manually inspect each file for PII.
Why it's wrong here
Manual inspection does not scale and is not maintainable for a large corpus. Converting to Parquet does not remove PII or normalize whitespace. This approach is error-prone, lacks repeatability, and would not preserve document boundaries in a structured way. It also introduces additional storage overhead without providing automated cleaning.
- ✗
Use the ai_classify function to label each document as 'clean' or 'dirty' and drop dirty ones.
Why it's wrong here
ai_classify assigns custom labels based on provided categories, but it does not actually remove PII or normalize text. Dropping 'dirty' documents would discard potentially valuable data rather than cleaning it. This approach does not satisfy the requirement to preserve document boundaries while cleaning the content, and it relies on subjective classification.
- ✗
Use the ai_analyze_sentiment function to filter out negative documents before training.
Why it's wrong here
Sentiment analysis is unrelated to PII removal or whitespace normalization. Filtering by sentiment would bias the training data and is not a data cleaning step for this scenario. The requirement is to clean PII and normalize whitespace, not to select documents based on emotional tone. This action would reduce dataset diversity without addressing the stated needs.
- ✓
Apply a regular expression to replace all whitespace sequences with a single space in the raw_text column.
Why this is correct
Normalizing whitespace with a regular expression such as regexp_replace(raw_text, '\\s+', ' ') collapses multiple spaces, tabs, and newlines into single spaces. This standardizes the text for tokenization and reduces noise without altering document boundaries. It is a scalable and deterministic step that can be applied in a SELECT or withColumn transformation.
- ✓
Use the ai_mask function to redact PII entities from the raw_text column.
Why this is correct
ai_mask is a built-in Databricks SQL function that detects and masks specified PII entities (such as names, emails, phone numbers) in text. It operates at scale within SQL or PySpark, making it suitable for large corpora. Applying it to raw_text removes PII while preserving the rest of the document content, which is essential for compliant fine-tuning data.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.