Databricks-ML-Assoc Model Development Practice Question
A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?
⚠ Common exam trap
Candidates often default to One-Hot Encoding, ignoring that high-cardinality categorical features cause dimensionality explosion, making Target Encoding the more scalable and performant choice for linear models.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Target Encoding to map categories to their mean target value.
Linear models require numeric inputs and are sensitive to the scale and representation of categorical variables. One-hot encoding high-cardinality features results in sparse, high-dimensional matrices that degrade performance. Target encoding or using specialized embeddings is more efficient. Choosing the right encoding strategy is a critical model development decision that directly impacts training speed, convergence, and the overall predictive accuracy of the final model deployed into production.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encode all categorical variables.
Why it's wrong here
One-hot encoding leads to the 'curse of dimensionality' when dealing with high-cardinality features. This creates an extremely sparse feature space that requires significant memory and compute, often causing linear models to overfit or run into memory constraints during the training phase, especially on large datasets.
- ✓
Use Target Encoding to map categories to their mean target value.
Why this is correct
Target encoding is highly effective for high-cardinality features, as it transforms categorical levels into a continuous numerical representation based on the target variable. This reduces dimensionality while retaining meaningful information, which is ideal for linear models that perform best when features are appropriately scaled and numeric.
- ✗
Drop the categorical features entirely.
Why it's wrong here
Dropping high-cardinality features frequently leads to significant information loss. If the categorical variable contains predictive power, removing it will negatively impact the model's accuracy. A better approach is to leverage encoding techniques that transform these variables into usable numeric representations rather than discarding the data prematurely.
- ✗
Convert all features into a single string representation.
Why it's wrong here
Combining features into a single string does not provide a usable format for linear models, which require numeric input vectors for matrix multiplication. This approach essentially creates a non-interpretable feature that the model will be unable to learn from, resulting in essentially random or poor predictions.
About these practice questions
One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.