Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: During data preparation for a regression model, a…

During data preparation for a regression model, a data scientist notices that two features have a Pearson correlation coefficient of 0.95. The scientist is concerned about multicollinearity. Which action should be taken to address this issue?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Remove one of the two features

Removing one of the highly correlated features reduces multicollinearity without losing much information, as they are nearly linearly dependent. Standardization does not fix multicollinearity. PCA would reduce dimensionality but may harm interpretability. Keeping both can destabilize coefficient estimates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply PCA to reduce dimensionality to a single component

    Why it's wrong here

    PCA produces linear combinations of both features, destroying their individual interpretability and discarding the variance each contributes separately, which regression coefficients then cannot express. PCA suits unsupervised dimensionality reduction on many correlated predictors, not a targeted fix for exactly two correlated features.

  • ✓

    Remove one of the two features

    Why this is correct

    A Pearson coefficient of 0.95 between two features signals strong linear redundancy, so their coefficients become unstable and interpretation unreliable. Removing one feature eliminates the collinearity directly, satisfying the stem's multicollinearity concern while retaining the remaining feature's predictive information.

  • ✗

    Keep both features as they are because linear models are robust to multicollinearity

    Why it's wrong here

    Ordinary least squares regression assumes no perfect multicollinearity; at 0.95 the coefficient estimates become unstable with inflated variance, so retaining both features harms the model. Keeping correlated features is defensible for tree-based models, which split on individual features and tolerate correlation.

  • ✗

    Standardize both features using StandardScaler

    Why it's wrong here

    StandardScaler only rescales each feature to zero mean and unit variance; it leaves the 0.95 correlation between them completely unchanged, so the multicollinearity persists. Standardisation is correct when features differ wildly in scale and a distance- or gradient-based algorithm needs comparable magnitudes.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.