MLA-C01 ML Model Development Practice Question
A financial services company is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.2% fraudulent transactions. The team wants to optimize the model for recall at a fixed precision of 90%. They are using the SageMaker built-in XGBoost algorithm with binary:logistic objective. Which evaluation metric should they monitor during training and hyperparameter tuning?
⚠ Common exam trap
The trap here is defaulting to ROC-AUC because it is common, but ROC-AUC can be overly optimistic on imbalanced data and does not enforce a precision constraint.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Precision-Recall AUC (PR-AUC)
For imbalanced classification with a precision constraint, the precision-recall curve is the right tool. PR-AUC summarizes performance across thresholds and emphasizes the minority class. The team can use the PR curve to find the threshold where precision is 90%, then read off recall. This directly aligns with their goal of maximizing recall at that precision, unlike ROC-AUC or single-threshold metrics such as F1 or MCC.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
F1 score
Why it's wrong here
F1 score is the harmonic mean of precision and recall, balancing both equally. However, it does not allow targeting a specific precision level such as 90%. The team's requirement is to maximize recall while maintaining precision at 90%, which F1 does not directly encode. F1 could be used as a general metric, but it does not satisfy the constraint.
- ✓
Precision-Recall AUC (PR-AUC)
Why this is correct
PR-AUC summarizes the precision-recall curve, which is more informative than ROC for imbalanced datasets. It captures the trade-off between precision and recall, allowing the team to select a threshold that achieves 90% precision and then measure recall. By monitoring PR-AUC, they can compare models' ability to rank positive instances highly, which directly supports optimizing recall at a fixed precision.
- ✗
Area Under the ROC Curve (AUC)
Why it's wrong here
AUC measures the trade-off between true positive rate and false positive rate across all thresholds, but it does not directly reflect performance at a specific precision. In highly imbalanced datasets, AUC can be misleadingly high even when precision at a desired recall is poor. The team needs a metric tied to a fixed precision, so AUC is not the best choice here.
- ✗
Matthews Correlation Coefficient (MCC)
Why it's wrong here
MCC considers all four confusion matrix categories and is robust to imbalance, but it produces a single value that does not allow targeting a specific precision. The team needs to fix precision at 90% and maximize recall, which requires a threshold-dependent analysis. MCC does not provide the granularity to enforce that constraint, making it less suitable for this scenario.
Go deeper
Related to this question
About these practice questions
Courseiva writes every MLA-C01 question from scratch — 665 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.