AI0-001 AI Models and Data Engineering Practice Question
Which THREE are common causes of data leakage in machine learning pipelines?
⚠ Common exam trap
CompTIA often tests the distinction between valid data splitting practices and actual leakage causes, so candidates may incorrectly select time-based splitting (Option A) as a leakage cause when it is actually a proper technique for sequential data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Using future information to predict the present
Option B is correct because using future information to predict the present is the classic definition of temporal leakage: when training features contain values that would not have been available at prediction time (e.g., tomorrow's price used to predict today's), the model learns relationships that cannot generalize. Option D is correct because applying normalization (e.g., StandardScaler or MinMaxScaler) before splitting lets the scaler compute statistics such as the mean, variance, or min/max over the test data, so test-set information leaks into the training transformation; the scaler must be fit only on the training split. Option E is correct because features directly derived from the target variable (e.g., a 'remaining balance' column computed from the label, or target-encoded aggregates that include the current row's label) encode the answer into the inputs, producing artificially high validation scores that collapse in production. Option A is not a leakage cause but a mitigation: time-based splitting is the recommended approach for sequential data precisely to prevent temporal leakage. Option C is also not inherently leakage: cross-validation on the entire dataset is standard practice as long as the preprocessing is fit within each fold; leakage arises only if transformations are fit on the full dataset before cross-validation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Using time-based splitting for sequential data
Why it's wrong here
Time-based splitting preserves chronological order so future data never informs past predictions; it prevents leakage rather than causing it. It is the correct choice for sequential or temporal data, where random splitting would let the model train on observations that occur after the validation period.
- ✓
Using future information to predict the present
Why this is correct
Using future information to predict the present leaks target-correlated data backwards through time. In temporal pipelines, features computed from later events encode outcomes unavailable at prediction time, inflating validation scores while the deployed model cannot access that information.
- ✗
Using cross-validation on the entire dataset
Why it's wrong here
Fitting preprocessing or feature selection before splitting lets validation folds influence training, but the leakage arises from ordering, not from cross-validation itself. Cross-validation is the correct tool for estimating generalisation error on i.i.d. data once splitting and fitting are confined to each training fold.
- ✓
Applying normalization before splitting data into train and test sets
Why this is correct
Fitting normalisation parameters on the full dataset before splitting lets test-set statistics influence training. The scaler's mean and standard deviation encode information from the held-out rows, so the model indirectly sees test data, producing optimistic evaluation scores.
- ✓
Including features that are directly derived from the target variable
Why this is correct
Features derived directly from the target encode the label itself, so the model learns a trivial mapping rather than genuine predictive signal. This is classic leakage: the feature would not exist at inference time, yet it makes training performance unrealistically high.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.