DA0-002 Data Analysis Practice Question
A data analyst is working with a dataset that includes a categorical variable 'education_level' with four categories: High School, Bachelor's, Master's, and PhD. The analyst wants to include this variable in a linear regression model. Which encoding method should the analyst use to avoid the dummy variable trap?
⚠ Common exam trap
The trap here is thinking that one-hot encoding all categories is fine; it actually creates perfect multicollinearity and unstable estimates.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
One-hot encoding with three categories (dropping one)
The dummy variable trap occurs when all dummy variables are included in a model with an intercept, causing perfect multicollinearity. To avoid it, one category is omitted as the reference. One-hot encoding with three categories (k-1) achieves this. Label encoding and binary encoding impose ordinality, which is inappropriate for nominal data. Including all four categories would cause multicollinearity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Label encoding
Why it's wrong here
Label encoding assigns arbitrary integers to categories (e.g., High School=1, Bachelor's=2, etc.), which implies an ordinal relationship that may not exist. This can mislead the regression model into treating the categories as having a natural order. Moreover, label encoding does not avoid the dummy variable trap; it simply uses one column but introduces incorrect assumptions.
- ✗
Binary encoding
Why it's wrong here
Binary encoding converts categories into binary code (e.g., 00, 01, 10, 11), reducing dimensionality but creating a false ordinal relationship and complex interactions. It is not designed to avoid the dummy variable trap and can still introduce multicollinearity if not used carefully. For linear regression, it is less interpretable than one-hot encoding with a reference category.
- ✓
One-hot encoding with three categories (dropping one)
Why this is correct
One-hot encoding with k-1 categories (here, three) avoids the dummy variable trap by preventing perfect multicollinearity. The dropped category becomes the reference level, and the coefficients for the other categories represent the difference from that reference. This is the standard approach for including nominal categorical variables in linear regression.
- ✗
One-hot encoding with all four categories
Why it's wrong here
One-hot encoding with all four categories creates four binary columns, which leads to perfect multicollinearity (the dummy variable trap) because the sum of all four columns equals 1 (the intercept). This makes the design matrix singular, causing unstable coefficient estimates. To avoid the trap, one category must be dropped.
Go deeper
Related to this question
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.