AI0-001 AI Models and Data Engineering Practice Question
Exhibit
The following output is from an MLflow run: Run ID: abc123 experiment_id: 1 status: FINISHED start_time: 2023-10-01 10:00:00 end_time: 2023-10-01 10:05:00 params: learning_rate: 0.01 max_depth: 10 n_estimators: 100 metrics: train_accuracy: 0.999 val_accuracy: 0.82 val_f1: 0.79 tags: model_type: RandomForest dataset: churn_v2
Refer to the exhibit. A data scientist reviews the MLflow run for a Random Forest model on customer churn data. What is the most likely issue with this model?
⚠ Common exam trap
CompTIA often tests the misconception that high validation accuracy alone indicates a good model, ignoring the critical comparison between training and validation metrics to detect overfitting.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The model is overfitting because there is a large gap between train and validation accuracy.
A large gap between training accuracy (e.g., 0.99) and validation accuracy (e.g., 0.82) indicates that the Random Forest model has memorized the training data but fails to generalize to unseen validation data. This is the classic symptom of overfitting, where the model captures noise rather than the underlying pattern. In MLflow, comparing train and validation metrics directly reveals this discrepancy.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model is underfitting because training accuracy is too high.
Why it's wrong here
Underfitting means both training and validation accuracy are low; high training accuracy indicates the opposite, overfitting, so the stated cause contradicts the symptom. It is tempting because high training accuracy is genuinely a warning sign, and would be correct if the option instead attributed the gap to overfitting.
- ✓
The model is overfitting because there is a large gap between train and validation accuracy.
Why this is correct
A wide gap between training and validation accuracy is the diagnostic signature of high variance: the Random Forest has memorised training noise rather than generalising. That gap, not the absolute accuracy, is what identifies overfitting in the MLflow run for this churn model.
- ✗
The model is performing well because validation accuracy is above 0.8.
Why it's wrong here
Validation accuracy above 0.8 does not establish good performance when the exhibit's training and validation metrics diverge or the dataset is imbalanced. It is tempting because accuracy thresholds are commonly cited as success criteria, and would be correct if both curves tracked closely on a representative, balanced holdout set.
- ✗
The model has a data leak because dataset version is v2.
Why it's wrong here
A dataset version number alone does not indicate leakage; leakage arises from target-derived features or improper train/test splitting, which the version label cannot reveal. It is tempting because versioning suggests reproducibility concerns, and would be correct if the exhibit showed preprocessing fitted on the full dataset before splitting.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.