A data scientist is preparing a dataset for training a classification model. The dataset contains 10,000 records with a binary target variable where 9,500 belong to class A and 500 belong to class B. Which technique should the scientist use to address the class imbalance?
Trap 1: Random undersampling of class A
Undersampling reduces data and may lose important patterns.
Trap 2: Adding Gaussian noise to class B
Adding noise does not create new informative samples.
Trap 3: Principal Component Analysis (PCA)
PCA reduces features, not address imbalance.
- A
SMOTE (Synthetic Minority Oversampling Technique)
SMOTE generates synthetic minority-class samples by interpolating between existing class B neighbours, rebalancing the 9,500-to-500 skew so the classifier no longer favours class A. This directly addresses the stem's imbalance constraint, unlike random undersampling, which would discard valuable majority records.
- B
Random undersampling of class A
Why it fails: Undersampling reduces data and may lose important patterns.
- C
Adding Gaussian noise to class B
Why it fails: Adding noise does not create new informative samples.
- D
Principal Component Analysis (PCA)
Why it fails: PCA reduces features, not address imbalance.