Courseiva

AI0-001 Machine Learning and Deep Learning Practice Question

A healthcare organization wants to use patient data to predict disease risk. They are concerned about bias in the model. Which step is most critical during the data preparation phase to mitigate bias?

⚠ Common exam trap

CompTIA often tests the misconception that bias can be fixed by technical tweaks like oversampling or removing sensitive attributes, when in fact the root cause is almost always unrepresentative training data that must be addressed at the collection or sampling stage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Ensuring the training data is representative of the target population

Ensuring the training data is representative of the target population is the most critical step during data preparation to mitigate bias because bias often originates from skewed or incomplete data that does not reflect the real-world distribution of patient demographics, conditions, and outcomes. Without a representative dataset, any subsequent preprocessing or modeling will propagate and potentially amplify existing disparities, leading to unfair or inaccurate predictions for underrepresented groups.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Applying SMOTE to oversample minority classes

    Why it's wrong here

    SMOTE balances class distribution, not representation bias arising from unrepresentative sampling or missing demographic groups. It is tempting because it genuinely helps when a minority class is under-represented and the model ignores it, but oversampling synthetic points cannot correct skewed or absent patient data.

  • ✗

    Using a more complex algorithm

    Why it's wrong here

    Algorithm complexity affects model capacity, not the representativeness of the training data; bias entering during collection or labelling persists regardless. It is tempting when accuracy metrics look poor, and would be relevant for improving predictive performance rather than fairness.

  • ✗

    Removing all demographic features

    Why it's wrong here

    Dropping demographic features does not remove bias, since correlated proxies such as postcode or utilisation patterns still encode it, and it blocks auditing for disparate impact. It is tempting as a naive fairness fix; it would be correct only where those attributes are legally prohibited from processing.

  • ✓

    Ensuring the training data is representative of the target population

    Why this is correct

    Bias originates in data whose composition misrepresents the population the model will serve. Ensuring the training set reflects the target population's demographics and disease prevalence prevents systematic underrepresentation, directly satisfying the data-preparation constraint before any modelling begins.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.