Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A data analyst is preparing a dataset for a machine learning model. The dataset contains a 'country' column with 150 unique values. To reduce dimensionality, the analyst wants to group less frequent countries into an 'Other' category. Which technique is being applied?

⚠ Common exam trap

Many exam-takers confuse binning with one-hot encoding, but one-hot encoding expands the feature space while binning reduces it.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Binning

The technique is binning, specifically grouping infrequent categories into an 'Other' bin. This reduces the cardinality of the 'country' feature, which can help prevent overfitting and improve model training efficiency. Binning is a common approach for handling high-cardinality categorical variables by consolidating rare levels into a single category.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding creates a binary column for each unique value, which would increase dimensionality rather than reduce it. With 150 countries, this would add 150 columns, worsening the curse of dimensionality. The analyst's goal is to reduce the number of categories, so one-hot encoding is the opposite of what is needed. It is a common technique but not applicable here.

  • ✗

    Normalization

    Why it's wrong here

    Normalization scales numerical values to a standard range, such as 0 to 1. It does not apply to categorical variables like country names. The analyst's task is to reduce the number of categories, not to rescale numeric magnitudes. Normalization would be irrelevant and could not be applied to a string column without prior encoding.

  • ✓

    Binning

    Why this is correct

    Binning (or bucketing) groups continuous or categorical values into a smaller number of bins. Here, grouping infrequent countries into 'Other' is a form of categorical binning. This reduces the number of unique categories, simplifies the feature, and can improve model performance by limiting noise from rare categories. It is a standard dimensionality reduction technique for high-cardinality categorical variables.

  • ✗

    Imputation

    Why it's wrong here

    Imputation fills missing values with estimated ones. The scenario does not mention missing data; it focuses on reducing the number of unique country values. Imputation would not address the high cardinality issue. It is a data cleaning technique for missingness, not for grouping categories.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.