Courseiva
hardMultiple Select

MLA-C01 Practice Question: A machine learning engineer is evaluating a…

A machine learning engineer is evaluating a binary classification model for detecting fraudulent transactions. The dataset is highly imbalanced, and the cost of false negatives (missing a fraud) is very high. Which two evaluation metrics should the engineer consider? (Choose two.)

⚠ Common exam trap

AWS often tests the misconception that accuracy is always a good metric, but in imbalanced classification problems, accuracy is a trap because a naive model can achieve high accuracy by simply predicting the majority class.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

F1-score

Recall (Option C) is correct because it measures the proportion of actual fraud cases that the model successfully identifies (TP / (TP + FN)), directly addressing the scenario's high cost of false negatives—maximizing recall minimizes missed frauds. F1-score (Option A) is correct because it is the harmonic mean of precision and recall (2 × (Precision × Recall) / (Precision + Recall)), providing a balanced single metric that remains informative on highly imbalanced datasets where a model could otherwise achieve high recall by flagging everything. Accuracy (Option B) is not appropriate because with severe class imbalance a trivial model predicting 'no fraud' for all transactions can score very high accuracy while catching zero frauds. Precision (Option D) is relevant but not among the two marked correct; it focuses on false positives rather than the costly false negatives and is already incorporated into the F1-score. Mean absolute error (Option E) is a regression metric and does not apply to binary classification evaluation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    F1-score

    Why this is correct

    F1-score is the harmonic mean of precision and recall, so it penalises models that sacrifice either measure. On a highly imbalanced fraud dataset with costly false negatives, it captures the precision-recall trade-off that accuracy would hide behind the dominant negative class.

  • ✗

    Accuracy

    Why it's wrong here

    Accuracy counts both classes equally, so a model predicting "not fraud" for every transaction scores highly on this imbalanced dataset while catching no fraud. Accuracy suits balanced datasets where false positives and false negatives carry similar cost.

  • ✓

    Recall

    Why this is correct

    Recall measures the proportion of actual frauds correctly identified, directly reflecting the cost constraint: high false-negative cost means missing frauds is expensive. Maximising recall minimises those misses, though it must be balanced against precision to avoid excessive false alarms.

  • ✗

    Precision

    Why it's wrong here

    Precision measures how many flagged transactions are genuinely fraudulent, targeting false positives; it ignores the false negatives the scenario prioritises. Precision is the correct focus when investigating flagged cases is costly and missing fraud is tolerable.

  • ✗

    Mean absolute error

    Why it's wrong here

    Mean absolute error measures regression residuals, not class labels, so it cannot express false negatives in a binary classifier. It is tempting because it quantifies average prediction error, and would be the right choice when evaluating a regression model predicting continuous values such as transaction amount.

About these practice questions

Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

4 more ways this is tested on MLA-C01

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A company is building a binary classifier for credit default prediction. The dataset is highly imbalanced (98% no default). They want to maximize recall for the minority class while maintaining reasonable precision. Which metric should be optimized during hyperparameter tuning?

medium
  • A.AUC-ROC
  • ✓ B.F1 score
  • C.Accuracy
  • D.Precision

Why B: The F1 score is the harmonic mean of precision and recall, making it ideal for imbalanced datasets where both false positives and false negatives matter. Optimizing F1 balances recall for the minority class (defaults) with precision, aligning with the goal of maximizing recall while maintaining reasonable precision.

Variation 2. A data scientist has trained a binary classification model for fraud detection. The dataset is highly imbalanced (99% non-fraud, 1% fraud). After evaluation, the model shows an accuracy of 99%, but the recall for fraud cases is only 10%. Which metric should the data scientist prioritize to improve the model's performance for fraud detection?

medium
  • A.Log loss
  • ✓ B.F1-score
  • C.Precision
  • D.Area under the ROC curve (AUC-ROC)

Why B: F1-score balances precision and recall, making it more informative than accuracy for imbalanced datasets. AUC-ROC is also used but F1 directly addresses the trade-off between false positives and false negatives. Precision alone does not capture recall, and Log loss does not directly indicate recall improvement.

Variation 3. A data scientist wants to evaluate the performance of a binary classification model. The dataset is highly imbalanced with only 5% positive class. Which metric should be used to evaluate the model?

easy
  • A.Accuracy
  • B.Mean Squared Error
  • C.R-squared
  • ✓ D.F1-score

Why D: In highly imbalanced datasets (e.g., only 5% positive class), accuracy is misleading because a model that predicts the majority class for all instances would achieve 95% accuracy without any predictive power. The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance by balancing false positives and false negatives. It is the standard metric for binary classification on imbalanced data in AWS SageMaker and other ML platforms.

Variation 4. A data scientist is training a binary classification model using imbalanced data where the positive class is only 1% of the dataset. The scientist wants to maximize the recall for the positive class while maintaining reasonable precision. Which evaluation metric is most appropriate to tune during model selection?

easy
  • A.Log loss
  • B.Area under the ROC curve (AUC)
  • ✓ C.F1 score
  • D.Accuracy

Why C: The F1 score is the harmonic mean of precision and recall, making it ideal for imbalanced datasets where the positive class is only 1%. By tuning the F1 score, the data scientist directly balances the trade-off between maximizing recall (capturing true positives) and maintaining reasonable precision (avoiding false positives), which aligns with the stated goal.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.