Databricks-ML-Pro Model Development Practice Question
A machine learning engineer is developing a custom MLflow Python model that requires a pre-processing step using a scikit-learn pipeline. They want to log the model such that it can be served with the pipeline included. Which approach should they take?
⚠ Common exam trap
The trap here is assuming that logging the scikit-learn pipeline alone is sufficient, when custom logic may require a pyfunc wrapper.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a Python class that inherits from `mlflow.pyfunc.PythonModel`, include the pipeline in the `predict` method, and log with `mlflow.pyfunc.log_model`.
For custom models that include pre-processing and a scikit-learn pipeline, the recommended approach is to create a custom `mlflow.pyfunc.PythonModel` subclass. The `predict` method can invoke the pipeline and any additional logic. Logging with `mlflow.pyfunc.log_model` packages the model and its dependencies for serving. This ensures the entire pipeline is included.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a Python class that inherits from `mlflow.pyfunc.PythonModel`, include the pipeline in the `predict` method, and log with `mlflow.pyfunc.log_model`.
Why this is correct
Subclassing `mlflow.pyfunc.PythonModel` allows encapsulating custom pre-processing and the scikit-learn pipeline. The `predict` method can call the pipeline and any additional logic. Logging with `mlflow.pyfunc.log_model` saves the model with its dependencies, enabling serving with the full pipeline. This is the standard approach for custom models.
- ✗
Log the pipeline as a separate model and use MLflow's multi-model serving to chain them.
Why it's wrong here
MLflow does not natively support chaining multiple models in a single serving endpoint without custom code. While multi-model serving exists, it is for serving multiple models independently, not for chaining pre-processing and prediction. The pyfunc approach is designed for this integration.
- ✗
Log the scikit-learn pipeline directly with `mlflow.sklearn.log_model`.
Why it's wrong here
Logging the pipeline directly works if the entire model is a scikit-learn pipeline. However, the scenario specifies a custom MLflow Python model that requires a pre-processing step. If the custom model includes additional logic beyond the pipeline, logging just the pipeline would omit that logic. The custom model approach is more appropriate.
- ✗
Use `mlflow.sklearn.log_model` and pass the pipeline as an artifact, then load it manually in a custom serving script.
Why it's wrong here
Passing the pipeline as an artifact and loading it manually requires custom serving code, which is not natively supported by MLflow model serving. The goal is to log the model so it can be served directly. The pyfunc approach integrates the pipeline into the model itself, avoiding manual loading.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.