A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. The feature table contains a column `transaction_time` that is a timestamp. After creating the training set with `create_training_set`, the resulting DataFrame includes `transaction_time` but the model training code fails because the timestamp is not accepted by the XGBoost trainer. What is the most likely cause and correct resolution?
`create_training_set` returns all columns from the feature table, including timestamps, without altering their types. XGBoost requires numeric or boolean inputs, so a raw timestamp causes a failure. The correct fix is to drop the timestamp or derive numeric features such as hour of day or day of week. This preserves feature lineage while making the data compatible with the trainer.
Why this answer
`create_training_set` preserves the original data types of feature columns, so a timestamp column remains a timestamp. XGBoost cannot handle timestamp types directly, causing the training failure. The data scientist must either drop the column or transform it into numeric features such as hour, day of week, or time since a reference point.
This ensures compatibility while retaining useful temporal signals.
Exam trap
The trap here is assuming that Databricks Feature Store automatically converts non-numeric columns like timestamps into numeric formats suitable for all model trainers.