Courseiva
Data Analysis →hardMultiple Choice

DA0-002 Data Analysis Practice Question

A data analyst is evaluating a binary classification model for loan default prediction. The model achieves 98% accuracy on the test set, but the analyst notices that only 2% of loans in the dataset actually defaulted. The analyst is concerned that accuracy is misleading. Which metric should the analyst examine to better assess the model's ability to identify actual defaulters?

⚠ Common exam trap

The trap here is assuming that high accuracy always indicates a good model, even when the data is imbalanced and the cost of missing the minority class is high.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Recall

In imbalanced classification, accuracy can be misleading because a model that always predicts the majority class achieves high accuracy but fails to identify the minority class. Recall focuses on the minority class by measuring how many actual positives were correctly identified. For loan default prediction, missing a defaulter is costly, so recall is a key metric. The other metrics are for regression or model fit, not classification performance.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Mean Absolute Error (MAE)

    Why it's wrong here

    MAE calculates the average absolute difference between predicted and actual values for continuous outcomes. Like RMSE, it is a regression metric and does not apply to classification problems. It would not help the analyst understand how well the model identifies loan defaults, as it does not account for the class imbalance or the binary nature of the target.

  • ✗

    Adjusted R-squared

    Why it's wrong here

    Adjusted R-squared is used in regression analysis to indicate how well the independent variables explain the variance in a continuous dependent variable, while penalizing for the number of predictors. It is not a classification metric and cannot evaluate a binary classifier's performance on imbalanced data. Thus, it is irrelevant to the scenario.

  • ✗

    Root Mean Squared Error (RMSE)

    Why it's wrong here

    RMSE measures the average magnitude of errors in continuous predictions, not the classification performance on imbalanced binary outcomes. It would be appropriate for regression tasks, such as predicting loan amount, but it does not directly quantify how well the model identifies defaulters. Using RMSE here would not address the analyst's concern about the minority class.

  • ✓

    Recall

    Why this is correct

    Recall, also called sensitivity or true positive rate, measures the proportion of actual defaulters that the model correctly identifies. With only 2% defaults, a model predicting 'no default' for all cases would achieve 98% accuracy but zero recall. Therefore, recall is the appropriate metric to assess the model's ability to catch actual defaulters, directly addressing the analyst's concern.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.