DA0-002 Data Analysis Practice Question
A healthcare analytics team is building a predictive model to identify patients at high risk of readmission within 30 days of discharge. The dataset includes 50,000 patient records with 200 features, including demographics, vital signs, lab results, and historical admissions. The target variable is binary (readmitted or not). The team uses a logistic regression model and achieves an AUC of 0.72 on the test set. However, the model's calibration is poor: for patients predicted to have a 70% risk, the actual readmission rate is only 40%. The team wants to improve calibration without significantly reducing discrimination (AUC). The data scientist suggests applying Platt scaling. However, the team lead is concerned that Platt scaling may reduce the model's ability to rank patients correctly. Which of the following is the best course of action?
⚠ Common exam trap
Many candidates think Platt scaling changes the model's ranking (AUC), but in reality it applies a monotonic transformation that preserves rank order, so discrimination is unaffected.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply Platt scaling on a held-out validation set to recalibrate the predicted probabilities without refitting the original model.
Platt scaling is a post-processing technique that fits a logistic regression model on the predicted probabilities from the original model using a held-out validation set. This recalibrates the probabilities without altering the ranking of patients (the AUC remains unchanged), directly addressing the poor calibration while preserving discrimination. Option C correctly describes this procedure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove poorly calibrated predictions by discarding all patients with predicted risk between 0.3 and 0.7.
Why it's wrong here
Discarding patients with predicted risk between 0.3 and 0.7 deletes the very cases needing recalibration and biases the test set, leaving the underlying probability estimates unchanged. It is tempting because restricting to confident predictions appears to remove miscalibrated outputs, but Platt scaling fits a logistic transform on held-out scores instead.
- ✗
Ignore calibration because AUC is the only metric that matters for readmission risk models.
Why it's wrong here
Calibration matters clinically because risk thresholds drive intervention decisions; a 70% prediction that is truly 40% misleads care teams. AUC only measures ranking, so ignoring calibration leaves the stated defect unresolved. It is tempting when discrimination is the sole reported metric, but probability reliability is the requirement here.
- ✓
Apply Platt scaling on a held-out validation set to recalibrate the predicted probabilities without refitting the original model.
Why this is correct
Platt scaling fits a logistic regression on the model's raw scores using a held-out validation set, correcting probability estimates while leaving the underlying model and its ranking untouched. Because it is monotonic, discrimination and AUC are preserved, satisfying the calibration goal without refitting.
- ✗
Switch to a random forest model, which inherently produces better-calibrated probabilities.
Why it's wrong here
Random forests produce votes averaged across trees, which are typically pushed toward 0 and 1 and are therefore often poorly calibrated, not inherently reliable. It is tempting because ensembles are assumed to output trustworthy probabilities, but Platt scaling directly fits a sigmoid to the existing logistic regression scores without changing the algorithm.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.