Databricks-ML-Pro ML Ops Practice Question
A data science team trains a scikit-learn model on a Databricks cluster and needs the same feature-engineering logic to run identically in a nightly batch scoring job and in a real-time Model Serving endpoint. They want a single artifact that encapsulates preprocessing and the estimator. Which approach should they use?
⚠ Common exam trap
The trap here is believing a model signature or container image reference can carry executable preprocessing, when only a pyfunc wrapper actually bundles that logic into the artifact.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log the model with MLflow using the sklearn flavor and a custom pyfunc wrapper that includes the preprocessing steps.
Wrapping preprocessing and the estimator in a single pyfunc model lets MLflow serialize both as one artifact. Batch scoring loads that artifact with the pyfunc loader, and Model Serving uses the same artifact, so feature transformations are guaranteed identical and training-serving skew is avoided without duplicating logic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Log the model with MLflow using the sklearn flavor and a custom pyfunc wrapper that includes the preprocessing steps.
Why this is correct
The pyfunc flavor wraps arbitrary Python logic, so preprocessing and the estimator travel as one artifact. Batch jobs load it with mlflow.pyfunc.load_model and the serving endpoint loads the same artifact, guaranteeing identical transformations. This is the standard Databricks pattern for eliminating training-serving skew when feature logic must be shared.
- ✗
Register the raw estimator in Unity Catalog and reimplement preprocessing separately in the batch job and endpoint code.
Why it's wrong here
Duplicating preprocessing in two code paths invites divergence: a change in one place silently skews the other. The scenario explicitly requires identical logic, and separate implementations cannot guarantee that. Reimplementing logic also increases maintenance burden and defeats the purpose of a single encapsulated artifact.
- ✗
Package the preprocessing into a custom container image and reference the image from the model version metadata.
Why it's wrong here
A container image can carry code and dependencies, but attaching it to a model version's metadata does not make MLflow execute that logic during scoring. Model Serving and batch jobs load the logged model artifact, not arbitrary images referenced in metadata. This approach also complicates batch usage on clusters that do not use the image.
- ✗
Log the model with the sklearn flavor and pass the preprocessing function name as a signature parameter.
Why it's wrong here
Model signatures describe input and output schema, not executable transformations. Naming a preprocessing function in the signature does not cause it to run at inference. The sklearn flavor alone only serializes the estimator, so preprocessing would still be missing during scoring, producing inconsistent results between batch and real time.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.