You are deploying a model using MLflow, and you want to log the custom pre-processing logic alongside the model so it is automatically applied during inference. How should you achieve this?
The pyfunc flavor is the standard way to package custom logic with models in MLflow. It allows users to define a wrapper that executes arbitrary code, including pre-processing, at inference time. This ensures that the model is self-contained, simplifying deployment and guaranteeing that input data is processed correctly.
Why this answer
MLflow allows you to log custom Python functions or entire pipelines as part of the model using the 'pyfunc' flavor. By wrapping your pre-processing steps and model into a custom pyfunc class and logging it, you encapsulate both the transformation and the prediction logic. This is a critical best practice for production deployments, as it ensures consistent pre-processing between training and inference, eliminating training-serving skew and simplifying the deployment process for the serving endpoint.
Exam trap
Test-takers often select generic artifact logging functions instead of the specialized pyfunc wrapper required to bundle custom pre-processing logic with the model.