Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A data analyst is working with a dataset that includes a categorical variable 'education_level' with four categories: High School, Bachelor's, Master's, and PhD. The analyst wants to include this variable in a linear regression model. Which encoding method should the analyst use to avoid the dummy variable trap?

⚠ Common exam trap

The trap here is thinking that one-hot encoding all categories is fine; it actually creates perfect multicollinearity and unstable estimates.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

One-hot encoding with three categories (dropping one)

The dummy variable trap occurs when all dummy variables are included in a model with an intercept, causing perfect multicollinearity. To avoid it, one category is omitted as the reference. One-hot encoding with three categories (k-1) achieves this. Label encoding and binary encoding impose ordinality, which is inappropriate for nominal data. Including all four categories would cause multicollinearity.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Label encoding

    Why it's wrong here

    Label encoding assigns arbitrary integers to categories (e.g., High School=1, Bachelor's=2, etc.), which implies an ordinal relationship that may not exist. This can mislead the regression model into treating the categories as having a natural order. Moreover, label encoding does not avoid the dummy variable trap; it simply uses one column but introduces incorrect assumptions.

  • ✗

    Binary encoding

    Why it's wrong here

    Binary encoding converts categories into binary code (e.g., 00, 01, 10, 11), reducing dimensionality but creating a false ordinal relationship and complex interactions. It is not designed to avoid the dummy variable trap and can still introduce multicollinearity if not used carefully. For linear regression, it is less interpretable than one-hot encoding with a reference category.

  • ✓

    One-hot encoding with three categories (dropping one)

    Why this is correct

    One-hot encoding with k-1 categories (here, three) avoids the dummy variable trap by preventing perfect multicollinearity. The dropped category becomes the reference level, and the coefficients for the other categories represent the difference from that reference. This is the standard approach for including nominal categorical variables in linear regression.

  • ✗

    One-hot encoding with all four categories

    Why it's wrong here

    One-hot encoding with all four categories creates four binary columns, which leads to perfect multicollinearity (the dummy variable trap) because the sum of all four columns equals 1 (the intercept). This makes the design matrix singular, causing unstable coefficient estimates. To avoid the trap, one category must be dropped.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.