DA0-002 Data Analysis Practice Question
A data analyst is working with a dataset containing house prices. After building a multiple linear regression model, the analyst observes that the model performs well on training data but poorly on validation data. Which technique is most appropriate to address this issue?
⚠ Common exam trap
CompTIA often tests the distinction between overfitting and underfitting, and candidates mistakenly choose polynomial transformation or adding features thinking they will improve fit, when in fact they increase model complexity and worsen overfitting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply L2 regularization (Ridge)
The model is overfitting the training data, as evidenced by high performance on training data but poor performance on validation data. L2 regularization (Ridge) adds a penalty term proportional to the square of the coefficients, which shrinks them and reduces model complexity, thereby improving generalization to unseen data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Decrease the training data size
Why it's wrong here
Reducing training data leaves the model with less signal and typically degrades validation performance further. Shrinking the training set is used to speed up experimentation or simulate data scarcity, not to fix a model that generalises poorly from overfitting.
- ✗
Use a polynomial transformation
Why it's wrong here
Polynomial transformation adds higher-order terms, increasing model capacity and variance, which worsens the train-validation gap. It is the correct choice when the model underfits and the relationship between predictors and price is genuinely non-linear, not when it overfits.
- ✗
Increase the number of features
Why it's wrong here
Adding features increases model capacity and variance, widening the train-validation gap rather than closing it. Feature expansion is correct when the model underfits because relevant predictors are missing, not when it already memorises the training set.
- ✓
Apply L2 regularization (Ridge)
Why this is correct
L2 regularization (Ridge) adds a penalty proportional to the squared magnitude of coefficients, shrinking them towards zero and reducing variance without eliminating features. This directly counters the overfitting causing the train-validation performance gap, satisfying the scenario's need to improve generalisation on unseen house price data.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.