Courseiva
AI Concepts and Techniques →mediumMultiple Select

AI0-001 AI Concepts and Techniques Practice Question

A data scientist is preparing a dataset for a binary classification model. The dataset has 1000 samples, with 800 positives and 200 negatives. To evaluate the model properly, which THREE steps should they take? (Select THREE)

⚠ Common exam trap

CompTIA often tests the misconception that removing minority samples or relying solely on accuracy is acceptable for imbalanced datasets, when in fact these approaches degrade model performance and evaluation validity.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a stratified train-test split to preserve class proportions

Option B is correct because a stratified train-test split preserves the 80/20 class ratio (800 positives, 200 negatives) in both the training and test sets, ensuring the evaluation reflects the true class distribution and avoids sampling bias. Option C is correct because SMOTE generates synthetic minority-class samples by interpolating between existing minority instances, balancing the training set so the classifier does not become biased toward the majority class. Option E is correct because with imbalanced data, accuracy is misleading; precision, recall, and F1-score reveal how well the model identifies the minority class and balances false positives against false negatives. Option A is wrong because deleting minority samples discards valuable information and worsens the class imbalance problem. Option D is wrong because a model predicting all positives would achieve 80% accuracy while completely failing on the minority class, making accuracy an unreliable metric here.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove the minority class samples to make the dataset balanced

    Why it's wrong here

    Deleting the 200 minority samples discards real data and biases the model toward the majority class, harming recall on negatives. It is tempting because balancing class counts is a legitimate technique, and would be correct via oversampling, undersampling or class weighting rather than outright removal of minority records.

  • ✓

    Use a stratified train-test split to preserve class proportions

    Why this is correct

    A stratified split preserves the 80:20 positive-to-negative ratio in both training and test sets, satisfying the requirement for proper evaluation. Without stratification, random splitting could yield test sets with skewed class proportions, distorting precision and recall.

  • ✓

    Apply SMOTE (Synthetic Minority Over-sampling Technique) to balance the training set

    Why this is correct

    SMOTE synthesises new minority-class examples by interpolating between existing negatives, balancing the training set's 800:200 skew. This prevents the classifier from favouring the majority class, though it must be applied only to training data, never the test set.

  • ✗

    Report only accuracy as the evaluation metric

    Why it's wrong here

    Accuracy alone is misleading on an 80/20 split: a model predicting all positives scores 80% while detecting zero negatives. It is tempting because accuracy is the default, intuitive metric, and would suffice only when classes are roughly balanced and false positives and false negatives carry equal cost.

  • ✓

    Use precision, recall, and F1-score for evaluation

    Why this is correct

    Precision, recall and F1-score suit this imbalanced 800:200 split because accuracy would be misleading: predicting all positives yields 80% while detecting no negatives. These metrics derive from the confusion matrix, exposing false positives and false negatives separately, so the minority class's performance is measured rather than masked by the majority.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.