hardMultiple Choice
MLA-C01 Practice Question: A data scientist observes that a linear…
A data scientist observes that a linear regression model has many irrelevant features. They want to perform feature selection to improve generalization. Which method combines feature selection with model training using a penalty that can shrink coefficients to zero?
⚠ Common exam trap
MLA-C01 often tests the L1 vs L2 distinction — candidates pick Ridge thinking any regularization performs feature selection, but only L1 (Lasso) zeroes coefficients.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Lasso regression
Lasso regression (L1 regularization) adds a penalty equal to the absolute value of the coefficients, which can shrink some coefficients exactly to zero — effectively performing feature selection during training. This makes it the method that combines feature selection with model fitting in a single step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Ridge regression
Why it's wrong here
Ridge regression applies an L2 penalty that shrinks coefficients toward zero but never sets them exactly to zero, so no features are eliminated. It is tempting because it is regularised linear regression, and it is correct when handling multicollinearity, not when you need sparse feature selection.
- ✓
Lasso regression
Why this is correct
Lasso applies L1 regularisation, which adds a penalty proportional to the absolute coefficient values. This drives irrelevant feature coefficients exactly to zero during training, combining selection with fitting and improving generalisation as the stem requires.
- ✗
Recursive Feature Elimination (RFE)
Why it's wrong here
RFE repeatedly fits a model and prunes the weakest features by ranking, but its penalty does not shrink coefficients to exactly zero during training. It is tempting because it does perform feature selection, and it is correct when you want wrapper-based ranking with a chosen estimator, not embedded regularisation.
- ✗
Principal Component Analysis (PCA)
Why it's wrong here
PCA transforms features into uncorrelated principal components, producing new axes rather than selecting original features or shrinking coefficients to zero. It is tempting because it reduces dimensionality, and it is correct for compression or decorrelation, not for identifying which original features to keep in a linear model.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.