Courseiva
hardMultiple Choice

MLA-C01 Practice Question: A machine learning engineer is performing feature…

A machine learning engineer is performing feature selection for a regression model with 200 features. The dataset has 10,000 samples. The engineer wants to remove irrelevant features while keeping those that have a strong non-linear relationship with the target. Which feature selection method is best suited for this requirement?

⚠ Common exam trap

The trap is assuming that correlation or linear-model-based selection is sufficient for all relationships — candidates may pick Pearson or Lasso without noticing the question explicitly requires capturing non-linear relationships, which only mutual information handles among the options.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Mutual information

Mutual information captures any kind of dependency between a feature and the target, including non-linear relationships, making it the best choice when the engineer wants to detect strong non-linear associations. Unlike correlation-based methods, mutual information does not assume a linear relationship and can identify features that Pearson or linear-model-based methods would miss. It is well-suited for the 200-feature, 10,000-sample regression scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Lasso regularization (L1)

    Why it's wrong here

    Lasso's L1 penalty shrinks coefficients under a linear model, so it selects on linear contribution and misses non-linear effects. It is tempting because it performs embedded selection efficiently on wide datasets, and would be correct if the target relationship were linear.

  • ✗

    Recursive feature elimination (RFE) with a linear model

    Why it's wrong here

    RFE with a linear estimator ranks features by linear coefficients, so it cannot capture non-linear dependence and may discard genuinely predictive features. It is tempting because RFE handles high-dimensional sets well, and would suit this data if relationships were linear rather than non-linear.

  • ✗

    Pearson correlation coefficient

    Why it's wrong here

    Pearson correlation measures only linear association, so features with strong non-linear relationships to the target score near zero and get dropped. It is tempting because it is fast and interpretable for 200 features, and would be correct if the requirement were linear relationships.

  • ✓

    Mutual information

    Why this is correct

    Mutual information captures arbitrary non-linear dependence between each feature and the target, unlike correlation-based filters that detect only linear relationships. With 10,000 samples across 200 features it remains computationally practical, satisfying the non-linear relevance requirement.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.