Courseiva
Data Preparation →mediumMultiple Select

Databricks-GenAI-Assoc Data Preparation Practice Question

A Generative AI engineer is preparing a Delta table of product reviews for a retrieval-augmented generation application. The reviews contain HTML tags, inconsistent casing, and occasional very long paragraphs that exceed the embedding model's context window. The engineer wants to clean and normalize the text before chunking and embedding. Which two actions should the engineer take to directly address the stated quality issues? (Choose two.)

⚠ Common exam trap

The trap here is treating generic data hygiene steps like casing changes or row filtering as fixes, when the stated defects are markup noise and context-window overflow.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Split long paragraphs into overlapping chunks sized to the embedding model's token limit.

The reviews need markup removed and length controlled before embedding. Stripping HTML and normalizing whitespace cleans the text, while chunking with overlap sized to the model token limit prevents truncation of long paragraphs. Uppercasing, dropping short reviews, and archiving raw HTML do not resolve the stated noise and overflow problems.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Drop every review shorter than 100 characters to reduce noise in the corpus.

    Why it's wrong here

    Short reviews may still be relevant and valid; discarding them removes potential retrieval evidence without solving HTML noise or context-window overflow. Filtering by length is a judgment call about corpus composition, not a fix for the stated preprocessing problems. It also risks deleting useful signal from the RAG index.

  • ✗

    Store the raw HTML in a separate column so it remains available for future parsing needs.

    Why it's wrong here

    Preserving raw HTML does not clean the text that will be embedded; the embedded field would still contain tags unless a separate cleaning step is applied. Keeping a raw copy is a reasonable audit practice, but it does not by itself address markup noise or paragraph length. The question asks for actions that directly resolve the quality issues.

  • ✓

    Split long paragraphs into overlapping chunks sized to the embedding model's token limit.

    Why this is correct

    The reviews contain paragraphs that exceed the model context window, so chunking with overlap ensures no content is silently truncated and context is preserved across boundaries. Sizing chunks to the token limit directly addresses the stated overflow problem. Overlap helps retrieval quality by preventing answers that straddle a boundary from being split apart.

  • ✓

    Strip HTML tags and normalize whitespace using a Spark transformation before chunking.

    Why this is correct

    Removing HTML tags and collapsing whitespace directly handles the markup noise and inconsistent formatting described in the reviews. Doing this in a Spark transformation keeps the operation distributed and ensures every downstream chunk contains clean prose. It is a prerequisite for meaningful chunk boundaries, since tags and stray whitespace can corrupt token counts and embeddings.

  • ✗

    Convert all review text to uppercase to standardize casing across the corpus.

    Why it's wrong here

    Uppercasing is a destructive normalization that removes case information useful for named entities and sentence boundaries, and it does not address HTML tags or length limits. Standardizing casing usually means lowercasing only when the embedding model expects it, but blanket uppercase harms tokenization quality. It leaves the actual quality issues untouched.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.