Courseiva
Data Preparation →hardMultiple Select

Databricks-GenAI-Assoc Data Preparation Practice Question

A team is preparing a large text corpus for embedding generation with a foundation model endpoint on Databricks. They must reduce token cost and improve retrieval quality before vectorization. Which two preprocessing steps should be applied to the raw text? (Choose two.)

⚠ Common exam trap

The trap here is assuming aggressive token reduction like stop-word removal or truncation always lowers cost, when it can degrade embedding quality without meaningful savings.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deduplicate near-identical passages and drop documents that fail a minimum content-length threshold.

Removing boilerplate and normalizing whitespace cuts wasted tokens while sharpening the text signal, and deduplicating plus dropping undersized fragments avoids paying to embed redundant or low-information content. Uppercasing, stop-word stripping, and blind truncation either damage semantics or discard content, so they do not improve cost or retrieval quality.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove stop words and punctuation from every passage before generating embeddings.

    Why it's wrong here

    Modern embedding models rely on full natural-language context, including function words and punctuation, to place passages in vector space. Stripping them can distort meaning and hurt retrieval, and because tokenizers compress common words efficiently, the token savings are small compared with the semantic damage, making this a poor trade for embedding quality.

  • ✓

    Deduplicate near-identical passages and drop documents that fail a minimum content-length threshold.

    Why this is correct

    Near-duplicate passages waste embedding calls and crowd the index with redundant vectors, so removing them lowers cost and improves result diversity. Dropping documents below a minimum length eliminates fragments and stubs that produce low-information embeddings, which raises overall retrieval precision without discarding meaningful content.

  • ✗

    Truncate every document to its first 512 characters to guarantee uniform chunk size.

    Why it's wrong here

    Blind truncation discards the majority of long documents, so retrieval can never surface information beyond the cut point. It also produces chunks that end mid-sentence, harming embedding coherence. Uniform size is better achieved by semantic chunking with overlap, not by cutting content away from the source.

  • ✓

    Normalize whitespace and strip boilerplate such as navigation menus, headers, and repeated legal footers.

    Why this is correct

    Boilerplate and irregular whitespace inflate token counts without carrying meaning, so removing them lowers embedding cost and sharpens the semantic signal in each chunk. Normalization also reduces spurious variation between otherwise identical passages, which improves similarity matching during retrieval and keeps chunk boundaries aligned with real content.

  • ✗

    Convert all text to uppercase before tokenization to standardize casing.

    Why it's wrong here

    Uppercasing destroys casing cues that tokenizers and models use to distinguish proper nouns, acronyms, and sentence structure, which can degrade embedding quality. It does not reduce token count meaningfully because most tokenizers already handle case variants, so it adds preprocessing risk without the cost or retrieval benefit the team needs.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.