Courseiva
Model Development →mediumMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

Which technique is most effective for handling high-cardinality categorical features when training a tree-based model on Databricks?

⚠ Common exam trap

Candidates frequently choose one-hot encoding, ignoring that it causes feature explosion and performance degradation with high-cardinality features in distributed tree-based models.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Target encoding, replacing labels with the mean of the target variable.

Target encoding or utilizing native categorical support in libraries like LightGBM or CatBoost is highly effective for high-cardinality features. These methods prevent the dimensionality explosion associated with one-hot encoding, which would otherwise lead to sparse, massive feature matrices that degrade training performance and memory efficiency. By handling categories internally, models maintain predictive power while remaining performant in distributed Databricks environments.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding for all categorical variables.

    Why it's wrong here

    One-hot encoding creates a new binary column for every unique category. With high cardinality, this results in extremely high-dimensional, sparse data that increases memory usage and slows down the training process significantly, often leading to poor model convergence and difficulty in interpreting the final feature importance results.

  • ✓

    Target encoding, replacing labels with the mean of the target variable.

    Why this is correct

    Target encoding maps categories to the average value of the target for that specific category. This drastically reduces the dimensionality compared to one-hot encoding. It is particularly effective for tree-based models, as it captures the relationship between the category and the target without creating an unmanageable number of features.

  • ✗

    Standardization using a StandardScaler.

    Why it's wrong here

    StandardScaler is designed for continuous, numerical features, not categorical ones. Applying it to categorical labels is mathematically incorrect and will result in non-interpretable feature representations. It does nothing to address the dimensionality issues of categorical features and should only be used for scaling numeric data to a standard range.

  • ✗

    Removing all categorical features entirely.

    Why it's wrong here

    Removing features is an extreme measure that usually leads to significant information loss and reduced model accuracy. Categorical features often contain critical predictive signals. Instead of deletion, practitioners should use encoding techniques that preserve information while keeping the data structure efficient for the training algorithm's computational requirements.

About these practice questions

Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.