mediumMultiple Choice
MLA-C01 Practice Question: During data preparation for a regression model, a…
During data preparation for a regression model, a data scientist notices that two features have a Pearson correlation coefficient of 0.95. The scientist is concerned about multicollinearity. Which action should be taken to address this issue?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Remove one of the two features
Removing one of the highly correlated features reduces multicollinearity without losing much information, as they are nearly linearly dependent. Standardization does not fix multicollinearity. PCA would reduce dimensionality but may harm interpretability. Keeping both can destabilize coefficient estimates.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Apply PCA to reduce dimensionality to a single component
Why it's wrong here
PCA produces linear combinations of both features, destroying their individual interpretability and discarding the variance each contributes separately, which regression coefficients then cannot express. PCA suits unsupervised dimensionality reduction on many correlated predictors, not a targeted fix for exactly two correlated features.
- ✓
Remove one of the two features
Why this is correct
A Pearson coefficient of 0.95 between two features signals strong linear redundancy, so their coefficients become unstable and interpretation unreliable. Removing one feature eliminates the collinearity directly, satisfying the stem's multicollinearity concern while retaining the remaining feature's predictive information.
- ✗
Keep both features as they are because linear models are robust to multicollinearity
Why it's wrong here
Ordinary least squares regression assumes no perfect multicollinearity; at 0.95 the coefficient estimates become unstable with inflated variance, so retaining both features harms the model. Keeping correlated features is defensible for tree-based models, which split on individual features and tolerate correlation.
- ✗
Standardize both features using StandardScaler
Why it's wrong here
StandardScaler only rescales each feature to zero mean and unit variance; it leaves the 0.95 correlation between them completely unchanged, so the multicollinearity persists. Standardisation is correct when features differ wildly in scale and a distance- or gradient-based algorithm needs comparable magnitudes.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.