Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A data scientist is training a binary…

A data scientist is training a binary classification model using Amazon SageMaker. The dataset has a severe class imbalance (95% negative, 5% positive). The model achieves 99% accuracy but fails to identify positive cases correctly. Which action should the data scientist take to improve the model's ability to detect positive cases?

⚠ Common exam trap

It's easy for candidates to think oversampling (SMOTE) or changing the model type is the primary fix, but the exam tests understanding that evaluation metrics and threshold tuning are critical for imbalanced classification, not just data preprocessing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the F1 score as the evaluation metric and adjust the classification threshold based on the precision-recall curve.

In a severely imbalanced dataset (95% negative, 5% positive), accuracy is misleading. The F1 score balances precision and recall, and adjusting the classification threshold based on the precision-recall curve allows the model to prioritize recall for the minority class, directly improving detection of positive cases. This approach is recommended in SageMaker when using built-in algorithms or custom models with imbalanced data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch to a logistic regression model with balanced class weights.

    Why it's wrong here

    Balanced class weights adjust the loss function but do not change the underlying algorithm's linear decision boundary, which may still miss the 5% positive class. It is tempting because weighting is a standard imbalance remedy, yet the scenario needs a metric and sampling strategy targeting recall, not just a model swap.

  • ✗

    Use accuracy as the evaluation metric and retrain the model.

    Why it's wrong here

    Accuracy is misleading under 95/5 imbalance, since predicting all negatives scores 95%; retraining with the same metric cannot improve positive-case recall. It is tempting because accuracy is the default metric, but precision, recall or F1 are needed to expose the minority-class failures.

  • ✗

    Apply SMOTE (Synthetic Minority Over-sampling Technique) to the training data.

    Why it's wrong here

    SMOTE synthesises minority-class examples in feature space, which suits tabular data but not the stem's need to reweight the existing 5% positives; it also risks noisy synthetic points. It is tempting because it directly targets imbalance, and would be correct for a numeric dataset with enough minority samples to interpolate between.

  • ✓

    Use the F1 score as the evaluation metric and adjust the classification threshold based on the precision-recall curve.

    Why this is correct

    Accuracy is misleading under 95/5 imbalance, since predicting all negatives scores 95%. The F1 score balances precision and recall on the minority class, and moving the threshold along the precision-recall curve trades false positives for higher recall, directly improving positive-case detection.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.