Databricks-GenAI-Assoc Data Preparation Practice Question
A GenAI engineer is preparing a fine-tuning dataset from a Delta table in Unity Catalog that contains raw user feedback. The feedback text includes irregular capitalization, HTML tags, and excessive punctuation. The engineer needs to normalize the text using Spark NLP within a Databricks notebook, ensuring the pipeline is reproducible and scalable. Which approach should the engineer use to apply this transformation?
⚠ Common exam trap
The trap here is assuming that simple string functions or regex alone suffice for comprehensive text normalization, overlooking the need for specialized NLP annotators and pipeline reproducibility.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the spark-nlp library's DocumentAssembler and Normalizer annotators in a Spark ML Pipeline, and save the pipeline to MLflow.
The correct approach uses Spark NLP's DocumentAssembler and Normalizer within a Spark ML Pipeline, which provides distributed, reproducible text normalization. Saving the pipeline to MLflow ensures version control and reusability. This method is scalable and integrates well with Databricks, making it suitable for preparing large fine-tuning datasets with complex cleaning needs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Delta Live Tables to define a streaming pipeline that applies the lower() and trim() functions to the text column.
Why it's wrong here
Delta Live Tables are excellent for building reliable data pipelines, but using only lower() and trim() is insufficient for handling HTML tags and excessive punctuation. This approach oversimplifies the normalization requirements and does not leverage specialized NLP tools. It would not produce the clean text needed for high-quality fine-tuning data.
- ✓
Use the spark-nlp library's DocumentAssembler and Normalizer annotators in a Spark ML Pipeline, and save the pipeline to MLflow.
Why this is correct
Spark NLP provides DocumentAssembler and Normalizer annotators that can be combined in a Spark ML Pipeline, enabling scalable and reproducible text normalization. Saving the pipeline to MLflow ensures versioning and reproducibility. This approach leverages Spark's distributed processing and integrates with Databricks workflows, making it ideal for preparing large-scale fine-tuning datasets.
- ✗
Use Databricks SQL's regexp_replace function in a SELECT statement to clean the text, and create a new table with the cleaned data.
Why it's wrong here
While regexp_replace can perform simple text cleaning, it lacks the advanced normalization capabilities of Spark NLP, such as handling HTML tags and irregular punctuation systematically. It also does not provide a reproducible pipeline artifact. This approach may be sufficient for basic cleaning but is not as robust or scalable for complex GenAI data preparation.
- ✗
Use pandas UDFs with Python's re module to apply custom cleaning functions to each row, and cache the resulting DataFrame.
Why it's wrong here
Pandas UDFs can be used for custom text cleaning, but they introduce overhead and may not scale as efficiently as native Spark NLP annotators. Caching the DataFrame does not ensure reproducibility of the transformation logic. This method is less maintainable and lacks the built-in NLP capabilities needed for comprehensive text normalization.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.