Courseiva
Data Preparation →mediumMultiple Choice

Databricks-GenAI-Assoc Data Preparation Practice Question

When preparing data for a fine-tuning task, you realize the dataset is severely imbalanced. Which Databricks technique should you use to create a more balanced dataset?

⚠ Common exam trap

Candidates often suggest manual filtering or simple random sampling. Random sampling does not solve class imbalance, whereas stratified sampling specifically ensures minority classes are adequately represented in the training set.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the sampleBy() method to perform stratified sampling.

Using Spark's 'sample()' or 'sampleBy()' transformation allows you to perform stratified oversampling or undersampling to balance classes. This ensures that the fine-tuned model doesn't become biased toward the most frequent categories in the dataset. Proper balancing is essential for ensuring robust model performance across all target classes, preventing the model from underperforming on rare but critical edge cases in real-world applications.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Run a standard 'SELECT * FROM table' query.

    Why it's wrong here

    This query retrieves the data in its original, imbalanced state. It does nothing to correct the bias or address the representativeness of the dataset. For machine learning, you must actively manipulate the data distribution to ensure the model learns a balanced and accurate understanding of all relevant topics.

  • ✓

    Use the sampleBy() method to perform stratified sampling.

    Why this is correct

    Stratified sampling allows you to select specific proportions from each category, enabling effective oversampling of minority classes or undersampling of majority classes. This technique is standard in Spark for preparing balanced training datasets, ensuring that the model learns effectively from all classes, regardless of their original prevalence in the data.

  • ✗

    Increase the number of epochs during the training process.

    Why it's wrong here

    Increasing epochs will cause the model to overfit to the majority classes even more aggressively if the data is imbalanced. It does not address the underlying sampling problem at all. Training longer on imbalanced data is counterproductive and will likely decrease the model's accuracy on the minority classes.

  • ✗

    Use a UDF to delete all rows that belong to majority classes.

    Why it's wrong here

    Deleting data is a destructive approach that leads to significant data loss and potential bias in the other direction. It is far more professional to use undersampling or oversampling techniques to achieve the desired balance while retaining the maximum amount of information possible for the model to learn from.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.