Databricks-GenAI-Assoc Data Preparation Practice Question
A GenAI engineer is preparing a large text corpus for fine-tuning an LLM. The corpus contains many near-duplicate documents and documents in multiple languages. They need to reduce redundancy and ensure language consistency. Which two steps should be performed during data preparation? (Choose two.)
⚠ Common exam trap
Many exam-takers confuse tokenization or versioning features with actual data cleaning steps, leading to choices that do not remove duplicates or filter languages.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the language detection function from the spark-nlp library to filter documents to a single target language.
To reduce redundancy and ensure language consistency, the engineer should use MinHash with LSH to remove near-duplicate documents and apply language detection to filter to the target language. These steps directly address the stated goals and are scalable for large corpora.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the language detection function from the spark-nlp library to filter documents to a single target language.
Why this is correct
Language detection identifies the language of each document, allowing you to filter out documents that do not match the target language. This ensures consistency and prevents the model from learning from irrelevant languages, which is critical when fine-tuning for a specific language task.
- ✗
Train a custom tokenizer on the entire corpus to better handle multilingual text.
Why it's wrong here
Training a custom tokenizer can improve tokenization for specific domains, but it does not remove near-duplicates or filter languages. While it might help with multilingual text, it is not a data preparation step for deduplication or language consistency, and it requires significant computational resources.
- ✗
Use Delta Lake time travel to revert to a previous version of the dataset if duplicates are found.
Why it's wrong here
Delta Lake time travel is useful for auditing and rollback, but it does not actively remove duplicates or filter languages. It is a passive versioning feature, not a data cleaning step. Relying on it would not address the redundancy or language consistency issues.
- ✓
Compute MinHash signatures and apply locality-sensitive hashing to identify and remove near-duplicate documents.
Why this is correct
MinHash with LSH is a scalable technique for detecting near-duplicate documents by estimating Jaccard similarity. It is well-suited for large corpora because it reduces pairwise comparisons. Removing near-duplicates prevents the model from overfitting to repeated content and improves training efficiency.
- ✗
Apply a fixed-size chunking strategy to split all documents into 512-token segments before deduplication.
Why it's wrong here
Chunking is typically done after cleaning and deduplication. Applying fixed-size chunking before deduplication would fragment documents and make it harder to detect near-duplicates at the document level. It could also increase the volume of data to process unnecessarily.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.