Databricks-ML-Pro Model Development Practice Question
A data scientist is using MLflow to log a model that includes a custom preprocessing step. They want to ensure that the preprocessing is applied consistently during both training and inference. Which approach should they take?
⚠ Common exam trap
The trap here is assuming that logging preprocessing as an artifact or using a scikit-learn Pipeline always works; custom logic may require a PythonModel for full encapsulation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a custom `PythonModel` that includes both the preprocessing and the model, and log it with `mlflow.pyfunc.log_model`.
Creating a custom `PythonModel` that integrates preprocessing and prediction ensures that the entire pipeline is encapsulated in a single MLflow model. When logged, this model applies preprocessing automatically during inference, guaranteeing consistency. This is the most robust approach when preprocessing involves custom logic not easily represented in a standard pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use `mlflow.sklearn.log_model` and include the preprocessing steps in a scikit-learn `Pipeline` object.
Why it's wrong here
Using a scikit-learn `Pipeline` is valid if the preprocessing steps are scikit-learn transformers. However, if the preprocessing includes custom logic not expressible as a transformer, a `Pipeline` may not suffice. The question specifies a custom preprocessing step, which might require a `PythonModel` for full flexibility. Additionally, not all preprocessing can be encapsulated in a `Pipeline`.
- ✗
Log the preprocessing code as a separate artifact and manually apply it before calling the model during inference.
Why it's wrong here
Logging preprocessing as an artifact requires manual intervention during inference, which is error-prone and not scalable. It does not guarantee consistency because the inference pipeline must remember to apply the artifact. This approach lacks the automation and encapsulation provided by MLflow's model packaging.
- ✗
Log the model with `mlflow.pyfunc.log_model` and provide the preprocessing function as a separate file in the `artifacts` parameter.
Why it's wrong here
Providing the preprocessing function as an artifact does not automatically apply it during inference. The model's `predict` method would need to explicitly load and call that function, which requires custom code. This approach is similar to logging as an artifact and does not ensure consistent application without additional implementation.
- ✓
Create a custom `PythonModel` that includes both the preprocessing and the model, and log it with `mlflow.pyfunc.log_model`.
Why this is correct
A custom `PythonModel` can encapsulate both preprocessing and the model's prediction logic. When logged with `mlflow.pyfunc.log_model`, the entire pipeline is packaged as a single model. During inference, MLflow loads this model and applies the preprocessing automatically, ensuring consistency between training and inference.
About these practice questions
One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.