A data scientist trains a model on historical data and achieves high accuracy on both the training set and a held-out test set. However, when the model is deployed in production, it performs poorly on new, unseen data. Which issue is most likely the cause?
Data leakage is the correct diagnosis: it happens when the training features contain information that would not be available at inference time, such as the target itself, a future value, or a post-outcome field. The model exploits this hidden shortcut, achieving very high accuracy on both training and test splits during offline evaluation, but in real-world deployment the leaked signal is absent, causing immediate and drastic performance collapse. This pattern of artificially perfect historical evaluation followed by severe production failure is the classic signature of data leakage.
Why this answer
Data leakage occurs when information from outside the training dataset is inadvertently used to train the model, causing it to learn patterns that do not generalize to new data. In this scenario, the high accuracy on both training and test sets but poor production performance indicates that the test set was contaminated with information from the future or from the target variable, making the model appear accurate during validation but fail in real-world deployment.
Exam trap
The trap here is that candidates confuse high accuracy on both training and test sets with overfitting, but the key differentiator is that overfitting would show a significant gap between training and test accuracy, whereas data leakage produces deceptively high accuracy on both sets.
How to eliminate wrong answers
Option A is wrong because overfitting would show high training accuracy but low test accuracy, not high accuracy on both sets. Option B is wrong because underfitting would result in poor performance on both training and test sets, not high accuracy. Option D is wrong because concept drift refers to a change in the underlying data distribution over time after deployment, not a static mismatch between training and production data at the time of deployment.