AI0-001 AI Models and Data Engineering Practice Question
A fraud detection team trains a gradient boosted tree model on transaction data. During evaluation, the team notices the model performs extremely well on the training set but poorly on a holdout set drawn from the same time period. Investigation shows that a feature named 'chargeback_flag' is populated only after a dispute is resolved, sometimes weeks after the transaction. The team wants to deploy the model to score transactions in real time. Which action best addresses the problem?
⚠ Common exam trap
The trap here is treating high offline accuracy as a modeling problem to tune rather than recognizing that a post-outcome feature is leaking the label.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Remove the 'chargeback_flag' feature and retrain the model using only features that are available at transaction scoring time.
The 'chargeback_flag' is populated only after a dispute is resolved, which is after the transaction outcome is known, so it leaks the label into training. The correct fix is to identify features that are actually available at real-time scoring and retrain without the leaky field. Hyperparameter tuning, larger validation sets, or partial replacements do not remove the leak and will keep offline metrics misleadingly high.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the size of the holdout set so the evaluation becomes more statistically reliable.
Why it's wrong here
A larger holdout set reduces variance in the performance estimate but does not fix leakage. Because the flag is populated only after disputes resolve, the holdout set also contains the leaky feature, so the model will continue to score well offline while failing on live transactions where the flag has not yet been set.
- ✗
Replace the flag with a rolling average of the customer's past chargebacks to preserve some of its signal.
Why it's wrong here
A historical rolling average can be a legitimate feature, but it does not address the specific leak in this scenario, and if computed over the same window that includes the current dispute it can still leak future information. The team must first establish which features are available at scoring time and remove those that are not.
- ✓
Remove the 'chargeback_flag' feature and retrain the model using only features that are available at transaction scoring time.
Why this is correct
The flag is a label leak: it is recorded only after the outcome the model is supposed to predict is known, so it cannot exist when scoring a new transaction. Removing it and retraining on features that are genuinely available at inference time eliminates the leak and produces a model whose offline metrics better reflect real-time performance. This is the correct root-cause fix.
- ✗
Apply stronger regularization and reduce the model's maximum depth to prevent it from relying on the flag.
Why it's wrong here
Regularization constrains model complexity but does not remove a feature that directly encodes the target outcome. A tree can still split on the leaky flag with minimal depth, so the model continues to look excellent offline and fail in production. Tuning hyperparameters treats a symptom while leaving the data leakage in place.
Visual reference
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.