Courseiva
easyMultiple Choice

MLA-C01 Practice Question: A data scientist discovers that a dataset for…

A data scientist discovers that a dataset for binary classification contains 95% negative samples and 5% positive samples. Which technique is MOST appropriate to address the class imbalance?

⚠ Common exam trap

MLA-C01 often tests the misconception that any resampling works — candidates must distinguish SMOTE's synthetic interpolation from naive duplication (RandomOverSampler) and from destructive undersampling that throws away majority data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use SMOTE to generate synthetic samples for the minority class

SMOTE (Synthetic Minority Over-sampling Technique) generates new, synthetic minority-class samples by interpolating between existing minority neighbors, which increases the minority representation without simply duplicating rows. This helps the model learn the minority decision boundary and is the standard, most appropriate technique for a 95/5 imbalance. It preserves all majority information while enriching the minority class.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use SMOTE to generate synthetic samples for the minority class

    Why this is correct

    SMOTE generates synthetic minority-class samples by interpolating between existing positive instances, rebalancing the 95:5 distribution. This gives the classifier enough positive examples to learn the minority pattern, unlike plain resampling that discards data or duplicates points.

  • ✗

    Remove all samples from the majority class

    Why it's wrong here

    Deleting every majority-class sample discards the 95% of genuine negatives the model needs, leaving a single-class dataset that cannot train a binary classifier. It is tempting because undersampling the majority class is a recognised balancing technique, but that approach removes only a subset, preserving enough negatives to learn from.

  • ✗

    Increase the learning rate of the model

    Why it's wrong here

    Learning rate controls the size of gradient-descent weight updates; it does not change the 95:5 class ratio, so the model still favours the majority class. It is tempting because tuning learning rate does improve convergence, which is correct when training is unstable or slow, not for imbalance.

  • ✗

    Downsample the majority class to 5% of the original size

    Why it's wrong here

    Downsampling the majority class to 5% discards roughly 95% of negative records, destroying information and leaving a tiny dataset that biases the model. It is tempting because undersampling does balance class ratios, which suits very large datasets where majority examples are genuinely redundant.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.