Courseiva
Data Preparation for Machine LearninghardMultiple ChoiceObjective-mapped

MLA-C01 Data Preparation for Machine Learning Practice Question

A data science team is building a model to predict fraudulent transactions. The dataset has 1 million legitimate transactions and only 1,000 fraudulent ones. They plan to use Amazon SageMaker to train a model. Which data preparation technique should they apply to address the severe class imbalance before training?

⚠ Common exam trap

AWS often tests the misconception that simple random oversampling (Option B) is sufficient, but the trap is that it causes overfitting, whereas SMOTE's synthetic generation provides better generalization for imbalanced datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Use SMOTE (Synthetic Minority Oversampling Technique) to generate synthetic fraudulent samples.

SMOTE (Synthetic Minority Oversampling Technique) is the correct choice because it generates synthetic fraudulent samples by interpolating between existing minority class instances in feature space, rather than simply duplicating records. This creates more diverse and realistic training data, reducing overfitting risk while addressing the severe 1:1000 class imbalance. Amazon SageMaker's built-in algorithms and data processing capabilities can easily integrate SMOTE-applied datasets for training.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Apply data augmentation using image transformations because fraud detection is like image classification.

    Why it's wrong here

    This is a tabular dataset; image augmentation is not applicable.

  • Randomly oversample the fraudulent class to match the legitimate count by duplicating existing fraud records.

    Why it's wrong here

    Simple duplication risks overfitting and does not add new information.

  • Use SMOTE (Synthetic Minority Oversampling Technique) to generate synthetic fraudulent samples.

    Why this is correct

    SMOTE creates synthetic examples by interpolating between existing minority instances, reducing overfitting risk.

  • Randomly undersample the legitimate class to 1,000 samples to create a balanced dataset.

    Why it's wrong here

    Undersampling discards 999,000 legitimate samples, losing significant information.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 835 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.