Courseiva
Data Preparation →easyMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A data engineer needs to prepare a Delta table of customer reviews for embedding generation. The reviews contain HTML tags, inconsistent whitespace, and mixed casing that hurt embedding quality. Which preparation step should be applied before generating embeddings?

⚠ Common exam trap

The trap here is assuming that a storage or model configuration change can compensate for dirty input text, when the embedding model only sees the characters it is given.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Normalize the text by stripping HTML tags, collapsing whitespace, and applying consistent casing.

Embedding models encode the exact tokens they receive, so HTML markup, irregular whitespace, and inconsistent casing add noise that dilutes semantic signal. Normalizing the text before embedding removes that noise and produces vectors that better reflect meaning. Storage format and vector dimension do not change the input text, so they cannot solve the quality problem.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply a hash of the raw review text as the embedding input to guarantee deterministic vectors.

    Why it's wrong here

    Hashing produces a fixed-length representation that has no semantic relationship to the text meaning, so similar reviews would not produce similar vectors. Retrieval depends on semantic proximity, which a hash cannot provide. The raw text must be cleaned and passed to an embedding model, not hashed.

  • ✗

    Increase the embedding vector dimension to capture the additional characters in the raw text.

    Why it's wrong here

    Vector dimension is fixed by the chosen embedding model and cannot be tuned per corpus. Raising it would not help because the model was trained at a specific dimension, and the noise from HTML tags and inconsistent spacing would still be encoded. Cleaning the text addresses the actual problem instead of masking it.

  • ✗

    Store the reviews in a Parquet file instead of Delta to improve text compression before embedding.

    Why it's wrong here

    File format affects storage and read performance, not the content of the strings that reach the embedding model. Moving to Parquet leaves HTML tags, extra whitespace, and inconsistent casing untouched, so embedding quality does not improve. The preparation step must transform the text itself.

  • ✓

    Normalize the text by stripping HTML tags, collapsing whitespace, and applying consistent casing.

    Why this is correct

    Removing HTML tags, collapsing whitespace, and normalizing casing reduces noise that would otherwise dilute the embedding signal. Embedding models encode the literal tokens they receive, so markup and erratic spacing consume vector capacity without adding semantic value. Cleaning the text first yields more consistent and comparable vectors across the corpus.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.