Databricks-ML-Pro Model Development Practice Question
A data scientist is building a model that requires custom preprocessing logic that is not available in standard libraries. They need to ensure this logic is bundled with the model for inference. What is the recommended approach to encapsulate this custom logic?
⚠ Common exam trap
Candidates often mistakenly select generic deployment options or standard framework saving methods, forgetting that custom preprocessing logic requires the specialized MLflow PyFunc model flavor wrapper.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement a custom class inheriting from mlflow.pyfunc.PythonModel.
Using the MLflow PyFunc (Python Function) flavor is the standard way to package arbitrary logic with a model. It allows developers to define a custom wrapper that includes preprocessing, prediction, and post-processing steps. When the model is logged, the custom class and its environment dependencies are saved, ensuring that the exact same logic is executed during inference, regardless of the deployment target.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Save the preprocessing logic in a separate Python script.
Why it's wrong here
Keeping preprocessing in a separate script creates a maintenance burden and increases the risk of the model failing to find the code in production. Encapsulating the logic within the MLflow PyFunc wrapper ensures the preprocessing code travels with the model artifact, simplifying deployment and ensuring consistency.
- ✓
Implement a custom class inheriting from mlflow.pyfunc.PythonModel.
Why this is correct
Inheriting from PythonModel allows developers to define a custom 'predict' method. This method can include any necessary preprocessing or post-processing logic, ensuring that the model acts as a self-contained unit that receives raw data and produces the final output without requiring external code dependencies.
- ✗
Use a Spark UDF for preprocessing during inference.
Why it's wrong here
While Spark UDFs can be used for transformation, they are not part of the model artifact itself. Relying on an external UDF during inference leads to dependency management issues and potential mismatches between the training and production environments, whereas PyFunc ensures the logic is bundled within the model.
- ✗
Hardcode the preprocessing logic directly into the SQL query.
Why it's wrong here
Hardcoding preprocessing into SQL queries is rigid and error-prone, making it difficult to test or update the logic. It also limits the model's portability, as the logic is tied to the data layer rather than the model itself, preventing the model from functioning correctly in different data environments.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.