Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A data scientist observes that a linear…

A data scientist observes that a linear regression model has many irrelevant features. They want to perform feature selection to improve generalization. Which method combines feature selection with model training using a penalty that can shrink coefficients to zero?

⚠ Common exam trap

MLA-C01 often tests the L1 vs L2 distinction — candidates pick Ridge thinking any regularization performs feature selection, but only L1 (Lasso) zeroes coefficients.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Lasso regression

Lasso regression (L1 regularization) adds a penalty equal to the absolute value of the coefficients, which can shrink some coefficients exactly to zero — effectively performing feature selection during training. This makes it the method that combines feature selection with model fitting in a single step.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Ridge regression

    Why it's wrong here

    Ridge regression applies an L2 penalty that shrinks coefficients toward zero but never sets them exactly to zero, so no features are eliminated. It is tempting because it is regularised linear regression, and it is correct when handling multicollinearity, not when you need sparse feature selection.

  • ✓

    Lasso regression

    Why this is correct

    Lasso applies L1 regularisation, which adds a penalty proportional to the absolute coefficient values. This drives irrelevant feature coefficients exactly to zero during training, combining selection with fitting and improving generalisation as the stem requires.

  • ✗

    Recursive Feature Elimination (RFE)

    Why it's wrong here

    RFE repeatedly fits a model and prunes the weakest features by ranking, but its penalty does not shrink coefficients to exactly zero during training. It is tempting because it does perform feature selection, and it is correct when you want wrapper-based ranking with a chosen estimator, not embedded regularisation.

  • ✗

    Principal Component Analysis (PCA)

    Why it's wrong here

    PCA transforms features into uncorrelated principal components, producing new axes rather than selecting original features or shrinking coefficients to zero. It is tempting because it reduces dimensionality, and it is correct for compression or decorrelation, not for identifying which original features to keep in a linear model.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.