Courseiva
Data Preparation →mediumMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

A team is preparing data for a RAG system and needs to remove duplicates from a large collection of PDF text extracts. What is the most efficient way to perform de-duplication in Databricks?

⚠ Common exam trap

Candidates often attempt single-node Python loops or pandas-based string matching on massive text corpora, resulting in out-of-memory errors on large datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Execute a distributed 'dropDuplicates' operation in Spark.

Using Spark's 'dropDuplicates()' on the text content column or calculating a hash (e.g., MD5) of the text to identify duplicates is highly scalable. This is crucial because redundant information in the vector database can cause the retriever to prioritize duplicate entries, leading to biased results and inefficiency. Removing duplicates ensures that the search index remains focused and that the retrieved context is diverse and informative.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a Python 'set' in a local loop on the driver node.

    Why it's wrong here

    Performing set operations on the driver node for large-scale datasets will lead to an Out-of-Memory error. Databricks' strength is distributed processing, which this approach completely ignores. It is an inefficient and unreliable method that fails as soon as the data volume exceeds the driver's memory capacity.

  • ✓

    Execute a distributed 'dropDuplicates' operation in Spark.

    Why this is correct

    Spark's 'dropDuplicates' is optimized for large, distributed datasets. It efficiently identifies and removes duplicate rows across the entire cluster, making it the standard approach for large-scale de-duplication. This ensures high-quality training sets and efficient vector indices without requiring complex custom code for distributed data handling.

  • ✗

    Manually compare every document against every other document using nested loops.

    Why it's wrong here

    Nested loops result in O(n^2) complexity, which is computationally prohibitive for any meaningful dataset size. This will effectively hang the cluster and fail to complete. Efficient de-duplication requires algorithms that scale linearly or log-linearly, which nested loops fail to achieve, making this an extremely poor engineering choice.

  • ✗

    Ignore duplicates, as vector databases automatically handle them during insertion.

    Why it's wrong here

    Vector databases do not inherently remove duplicates; they store what they are given. Keeping duplicates wastes storage, increases query latency, and negatively impacts the quality of the retrieved context. It is a fundamental data quality issue that must be addressed during the data preparation stage, not ignored.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.