AI0-001 Implementing AI Solutions Practice Question
A data science team is preparing a dataset for a supervised learning task. They split the data into training and test sets. The team then normalizes the features using the mean and standard deviation calculated from the entire dataset before splitting. What issue does this introduce?
⚠ Common exam trap
AI0-001 often tests whether candidates recognize that preprocessing steps (scaling, imputation, encoding) must be fit only on training data — many candidates focus on model training and forget that leakage can occur in the data preparation pipeline.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It introduces train/test leakage
Computing normalization statistics (mean and standard deviation) on the entire dataset before splitting means the test set's statistics influence the training data transformation. This is a form of train/test leakage: information from the test set leaks into the training pipeline, producing an optimistically biased estimate of model performance. The correct approach is to fit the scaler on the training set only and apply the same transformation to the test set.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It improves model generalization
Why it's wrong here
This practice actually harms generalization due to leakage, not improves it.
- ✓
It introduces train/test leakage
Why this is correct
Computing the mean and standard deviation across the whole dataset lets statistics from the test set influence the scaling applied to training features. Test information therefore leaks into training, producing optimistic evaluation results that will not generalise to unseen data.
- ✗
It causes the model to overfit the training data
Why it's wrong here
Leakage of test statistics into training does not cause overfitting; it inflates reported test performance because the test set influenced the scaling parameters. Overfitting arises from excessive model capacity or too little regularisation, not from computing normalisation statistics across the full dataset.
- ✗
It reduces the variance of the features
Why it's wrong here
Standardisation rescales features to unit variance; it does not reduce the underlying variance of the data itself. The option is tempting because normalisation changes the numeric spread of values, yet the actual defect here is data leakage from computing statistics over the test set before splitting.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.