Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A data scientist is preparing a dataset for training a machine learning model. The dataset contains a mix of numerical and categorical features, and some features have high cardinality. The data scientist needs to apply appropriate encoding techniques to transform categorical variables into a format suitable for the model. Which TWO encoding methods are most appropriate for high-cardinality categorical features? (Choose two.)

⚠ Common exam trap

The trap here is assuming that one-hot encoding is always the default for categorical variables, without considering the dimensionality explosion with high-cardinality features.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Target encoding

Target encoding and frequency encoding are both effective for high-cardinality categorical features. Target encoding replaces categories with the mean target value, capturing predictive relationships, while frequency encoding replaces categories with their counts, reducing dimensionality without imposing order. One-hot encoding creates too many columns, label encoding imposes false order, and binary encoding may not capture target relationships as well. Therefore, target and frequency encoding are the most appropriate methods.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Target encoding

    Why this is correct

    Target encoding replaces each category with the mean of the target variable for that category. It is effective for high-cardinality features because it reduces dimensionality and captures the relationship between the category and the target. However, it can lead to overfitting if not properly regularized, especially with rare categories. Techniques like smoothing or adding noise can mitigate this. In this scenario, target encoding is a suitable method for high-cardinality categorical features.

  • ✗

    Label encoding

    Why it's wrong here

    Label encoding assigns a unique integer to each category. While it is memory-efficient, it imposes an ordinal relationship that may not exist, which can mislead models that interpret numerical values as ordered. For high-cardinality features, label encoding can still result in many unique integers, but the arbitrary ordering can introduce bias. It is generally not recommended for nominal categorical variables unless the model can handle categorical features natively. Thus, label encoding is not the best choice here.

  • ✓

    Frequency encoding

    Why this is correct

    Frequency encoding replaces each category with its frequency or count in the dataset. It is a simple and effective method for high-cardinality features because it reduces dimensionality and does not impose an artificial order. It can capture the importance of frequent categories, which may be informative. However, it does not consider the target variable, so it may not capture predictive relationships as well as target encoding. In this scenario, frequency encoding is a suitable choice for high-cardinality categorical features.

  • ✗

    Binary encoding

    Why it's wrong here

    Binary encoding converts each category into binary code and then splits the digits into separate columns. It reduces dimensionality compared to one-hot encoding and is more memory-efficient. However, it can still create many columns for very high cardinality and may not capture the relationship with the target as effectively as target encoding. It also introduces a somewhat arbitrary binary representation. While it is a valid technique, it is not as appropriate as target or frequency encoding for high-cardinality features in this context.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding creates a binary column for each category, which can lead to a very high-dimensional feature space when cardinality is high. This increases memory usage and computational cost, and can cause the curse of dimensionality, negatively impacting model performance. While it is simple and preserves all information, it is generally not recommended for high-cardinality features. Therefore, one-hot encoding is not the most appropriate choice in this scenario.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.