AI0-001 Machine Learning and Deep Learning Practice Question
A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?
⚠ Common exam trap
CompTIA often tests the misconception that bias can be fixed by technical tweaks like oversampling or removing sensitive attributes, when in fact the root cause is almost always unrepresentative training data that must be addressed at the collection or sampling stage.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Ensuring the training data is representative of the target population
Ensuring the training data is representative of the target population is the most critical step during data preparation to mitigate bias because bias often originates from skewed or incomplete data that does not reflect the real-world distribution of patient demographics, conditions, and outcomes. Without a representative dataset, any subsequent preprocessing or modeling will propagate and potentially amplify existing disparities, leading to unfair or inaccurate predictions for underrepresented groups.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Applying SMOTE to oversample minority classes
Why it's wrong here
SMOTE balances class distribution, not representation bias arising from unrepresentative sampling or missing demographic groups. It is tempting because it genuinely helps when a minority class is under-represented and the model ignores it, but oversampling synthetic points cannot correct skewed or absent patient data.
- ✗
Using a more complex algorithm
Why it's wrong here
Algorithm complexity affects model capacity, not the representativeness of the training data; bias entering during collection or labelling persists regardless. It is tempting when accuracy metrics look poor, and would be relevant for improving predictive performance rather than fairness.
- ✗
Removing all demographic features
Why it's wrong here
Dropping demographic features does not remove bias, since correlated proxies such as postcode or utilisation patterns still encode it, and it blocks auditing for disparate impact. It is tempting as a naive fairness fix; it would be correct only where those attributes are legally prohibited from processing.
- ✓
Ensuring the training data is representative of the target population
Why this is correct
Bias originates in data whose composition misrepresents the population the model will serve. Ensuring the training set reflects the target population's demographics and disease prevalence prevents systematic underrepresentation, directly satisfying the data-preparation constraint before any modelling begins.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.