A machine learning engineer is using Databricks AutoML to train a classification model. They notice that the best model from AutoML has a high F1 score on the validation set but performs poorly on a holdout test set. They suspect that the data has a temporal component and that the default train/validation split is causing data leakage. What should they do to address this?
Databricks AutoML supports a `time_col` parameter that specifies a time column for temporal data. When set, AutoML performs a chronological split, ensuring that training data precedes validation data. This prevents leakage from future data into the training set. It is the correct approach for time-series or temporally ordered data to get realistic validation performance.
Why this answer
Databricks AutoML provides the `time_col` parameter to handle temporal data. When set, AutoML performs a chronological split, ensuring that training data precedes validation data, which prevents leakage from future data. This is the correct way to address the poor holdout performance caused by a random split.
Other options do not fix the temporal leakage issue.
Exam trap
The trap here is assuming that increasing validation size or enabling cross-validation automatically respects temporal order, when only specifying the time column triggers a chronological split.