A data engineer is preparing a dataset in Amazon SageMaker Data Wrangler for a binary classification model. The dataset contains missing values in several numeric columns, and the engineer wants a reusable, reproducible transformation that can be applied identically to the training data and to future inference data. The engineer plans to export the transformation and integrate it into a SageMaker Pipeline. Which TWO actions should the engineer take to ensure the imputation is consistent between training and inference? (Choose two.)
Learning the fill statistic from the training data ensures the imputation parameters are derived from the correct distribution and not from inference data, which would leak information. Data Wrangler captures these learned parameters as part of the flow, making the transformation reproducible. This is the foundation for applying the same imputation consistently at inference time.
Why this answer
Consistency between training and inference requires that the imputation statistic be learned from training data and that the same transformation be reused at inference. Configuring Data Wrangler to learn the fill statistic from training data captures the parameters, and exporting the flow into a SageMaker Pipeline processing step ensures those exact parameters and logic are applied to both training and inference inputs.
Exam trap
The trap here is treating imputation as a one-time data cleanup rather than a parameterized transformation whose learned statistics must be persisted and reapplied at inference.