DA0-002 Data Analysis Practice Question
A data analyst at a retail company is building a multiple linear regression model to forecast weekly sales. The dataset contains 50 predictor variables, including store size, promotional spend, holiday indicators, and many others. After training the model, the analyst observes an R-squared of 0.99 on the training set but only 0.55 on the holdout test set. Which action should the analyst take first to address this discrepancy?
⚠ Common exam trap
Many candidates think a high R-squared is always good, or they may confuse overfitting with underfitting and choose to add more complexity (Option D) or more data (Option B), rather than recognizing the need to reduce model complexity and apply regularization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Remove highly correlated predictor variables and apply regularization (e.g., Ridge or Lasso).
The high R-squared of 0.99 on training data versus 0.55 on test data is a classic sign of overfitting, where the model has learned noise and specific patterns in the training set that do not generalize. Removing highly correlated predictors reduces multicollinearity and model complexity, while regularization (Ridge or Lasso) penalizes large coefficients, shrinking them to prevent overfitting. This is the most direct first step to improve generalization.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Remove highly correlated predictor variables and apply regularization (e.g., Ridge or Lasso).
Why this is correct
The 0.99 versus 0.55 gap indicates overfitting from 50 predictors, many correlated. Removing correlated variables reduces multicollinearity and dimensionality, while Ridge or Lasso regularisation penalises large coefficients, shrinking variance and improving holdout generalisation before other remedies are attempted.
- ✗
Add more predictor variables to increase the training R-squared further.
Why it's wrong here
Adding predictors worsens the problem: with 50 variables and a 0.99-versus-0.55 gap, the model is already overfitting, so extra features inflate training fit while holdout performance drops further. Adding variables suits underfitting, where both training and test scores are low.
- ✗
Use k-fold cross-validation with a different random seed to get a more reliable test set estimate.
Why it's wrong here
Reseeding k-fold changes only which rows land in each fold; the 0.99-to-0.55 gap persists because the model itself overfits 50 predictors. Cross-validation gives a steadier estimate, which is useful for comparing candidate models, not for closing a genuine generalisation gap.
- ✗
Increase the number of hidden layers in the model to capture more complexity.
Why it's wrong here
Hidden layers belong to neural networks, not linear regression; adding them replaces the model rather than addressing the discrepancy. Depth helps when a network underfits complex patterns, whereas here the training score of 0.99 against 0.55 signals overfitting that extra capacity would worsen.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.