Databricks-GenAI-Assoc Data Preparation Practice Question
A team is preparing a pretraining corpus and must filter out near-duplicate documents before tokenization. The corpus contains 500 million short text records in a Delta table. They want a scalable, deterministic deduplication signal that can be computed per record and compared across the dataset. Which approach best fits this requirement?
⚠ Common exam trap
The trap here is equating deduplication with exact-match hashing, when near-duplicate removal requires a similarity-preserving signature rather than a cryptographic digest.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compute a MinHash signature per document with a fixed number of hash permutations and compare signatures with a locality-sensitive hashing banding scheme.
Near-duplicate detection at scale needs a compact per-record signature plus a candidate-generation strategy, which MinHash with locality-sensitive hashing banding provides. Exact hashing misses near-duplicates, embedding clustering is costly and nondeterministic, and length-based filtering has no content signal, so none of those satisfy the deterministic, scalable requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Sort the corpus by document length and drop every record whose length falls within one standard deviation of another record.
Why it's wrong here
Length alone carries no content signal, so this removes many unique documents while retaining near-duplicates of similar length. It is also not a comparison of content across records in any meaningful way, making the result both lossy and ineffective for deduplication, and it fails the requirement for a deterministic similarity signal.
- ✗
Compute an MD5 hash of each full document and drop rows sharing the same hash value.
Why it's wrong here
An MD5 of the full text detects only exact duplicates; a single changed character produces a completely different digest. Near-duplicate documents in a pretraining corpus typically differ by boilerplate or whitespace, so this method would keep most of them and fail the stated goal of removing near-duplicates while still requiring a full shuffle to group hashes.
- ✓
Compute a MinHash signature per document with a fixed number of hash permutations and compare signatures with a locality-sensitive hashing banding scheme.
Why this is correct
MinHash with LSH banding produces a fixed-length signature per document and a scalable candidate-generation step, so near-duplicates are found without all-pairs comparison. It is deterministic given fixed seeds and permutations, distributes cleanly across Spark partitions, and directly targets near-duplicate detection rather than exact matches, fitting the 500 million record corpus.
- ✗
Train a sentence embedding model and cluster documents with k-means, then drop all but one document per cluster.
Why it's wrong here
Embedding plus k-means is expensive at 500 million records, requires choosing k, and is sensitive to initialization, so results are not deterministic across runs. Clusters also group semantically similar but legitimately distinct documents, causing false removals. It solves semantic clustering rather than the near-duplicate detection the team requested.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.