Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data scientist is preparing a dataset for training a binary classification model. The dataset has 100,000 rows and 50 features. The target variable is imbalanced, with only 5% positive cases. Which technique should the data scientist apply to address the class imbalance BEFORE training?

⚠ Common exam trap

AWS often tests whether candidates confuse data preprocessing techniques (scaling, encoding, dimensionality reduction) with methods that directly modify the class distribution, leading them to pick a plausible but irrelevant option like PCA or scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Random oversampling of the minority class

Random oversampling of the minority class (Option B) directly addresses the class imbalance by duplicating examples from the positive class until the class distribution is more balanced. This prevents the binary classification model from being biased toward the majority class, which is critical when only 5% of the 100,000 rows are positive cases. Oversampling is applied before training to ensure the model sees sufficient minority examples during learning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Principal Component Analysis (PCA) dimensionality reduction

    Why it's wrong here

    PCA reduces feature dimensionality by projecting onto principal components; it leaves the 5% positive prevalence untouched, so imbalance persists. PCA is appropriate for cutting redundant correlated features or speeding training, not for correcting skewed class distributions.

  • ✓

    Random oversampling of the minority class

    Why this is correct

    Random oversampling duplicates minority-class rows, raising the 5% positive rate toward parity so the algorithm no longer biases toward the majority class. It satisfies the pre-training constraint by rebalancing class distribution before the model sees the data, unlike threshold tuning applied afterwards.

  • ✗

    Standard scaling of numerical features

    Why it's wrong here

    Standard scaling normalises feature magnitudes; it leaves the 5% positive rate untouched, so the classifier still sees skewed priors. It is tempting because scaling is a genuine preprocessing step, but it belongs to feature preparation, not to resampling techniques such as class weights or SMOTE that alter the target distribution.

  • ✗

    One-hot encoding of categorical variables

    Why it's wrong here

    One-hot encoding converts categorical features into binary columns; it does not change the 5% positive rate, so the classifier still sees skewed classes. It is the correct preprocessing step when nominal features must be represented numerically, not when addressing class imbalance.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.