Courseiva
Data Analysis →mediumMultiple Choice

DA0-002 Data Analysis Practice Question

A healthcare analytics team is analyzing patient readmission rates. They have a dataset with thousands of records including patient age, diagnosis, length of stay, number of prior admissions, and discharge date. The goal is to identify key factors influencing readmission and create a model to predict high-risk patients. The data is imbalanced: only 5% of patients are readmitted within 30 days. The team plans to use logistic regression. What is the most appropriate approach?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply oversampling techniques like SMOTE to the training set

With imbalanced data, logistic regression can be biased toward the majority class. Oversampling the minority class (e.g., SMOTE) helps the model learn patterns for readmission. Using accuracy as a metric would be misleading. Removing majority samples discards valuable data. Using data as-is often fails to predict the minority class.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use the dataset as is because logistic regression handles imbalance

    Why it's wrong here

    Logistic regression does not correct class imbalance; with 5% positives it predicts the majority class and yields misleading accuracy. It is tempting because logistic regression is the standard binary classifier, and it would suffice if readmissions were closer to balanced or thresholds and weights were tuned.

  • ✗

    Remove most of the non-readmitted patients to balance the dataset

    Why it's wrong here

    Discarding most non-readmitted records throws away genuine information and biases the model, distorting the true 5% base rate. It is tempting because it equalises classes cheaply, and it would be defensible only when the majority class is contaminated or mislabelled rather than merely rare.

  • ✗

    Use accuracy as the evaluation metric

    Why it's wrong here

    With only 5% readmissions, accuracy is misleading: predicting "no readmission" for every patient scores 95% while identifying no high-risk cases. Accuracy suits balanced class distributions where each class contributes equally to the metric. Recall, precision or AUC-ROC expose the model's failure on the minority class.

  • ✓

    Apply oversampling techniques like SMOTE to the training set

    Why this is correct

    With only 5% readmissions, logistic regression would bias toward the majority class. SMOTE synthesises minority-class examples in the training set, balancing class distribution so the model learns readmission patterns rather than predicting 'no readmission' constantly.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.