Databricks-ML-Pro Model Deployment Practice Question
A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with maximum throughput and minimum latency. The model requires an external preprocessing Python script during inference. Which deployment approach best leverages MLflow and Databricks architecture?
⚠ Common exam trap
Candidates often suggest performing preprocessing in the Spark application before calling the model. This creates training-serving skew and latency issues compared to native PyFunc model packaging.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log the model and preprocessing code using custom pyfunc, register it in Unity Catalog, and configure a serverless Model Serving endpoint.
Packaging the custom preprocessing logic along with the scikit-learn estimator into an MLflow PyFunc model ensures that all inference data transformations happen inside the serving container natively. This prevents client-side processing bottlenecks and guarantees consistent feature engineering between training and real-time serving environments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Log the model and preprocessing code using custom pyfunc, register it in Unity Catalog, and configure a serverless Model Serving endpoint.
Why this is correct
Custom pyfunc models encapsulate both the estimator and arbitrary preprocessing code into the artifact. Registering this artifact in Unity Catalog enables secure governance, and deploying it to a serverless endpoint provides auto-scaling and low-latency inference capabilities.
- ✗
Deploy the raw scikit-learn model artifact to Model Serving and execute the preprocessing logic inside a separate Spark structured streaming job.
Why it's wrong here
Separating the preprocessing logic from the serving endpoint introduces significant network latency and operational complexity. Real-time REST endpoints require self-contained prediction pipelines to handle incoming single-record or micro-batch payloads efficiently without relying on external streaming jobs.
- ✗
Write the preprocessing code inside a Databricks notebook and call the model endpoint using a client-side REST API request.
Why it's wrong here
Relying on client-side preprocessing shifts the computational burden and synchronization risk to the calling application. This violates enterprise deployment best practices where the serving container should act as the single source of truth for the complete inference pipeline.
- ✗
Save the model weights to DBFS and build an external Flask application on Amazon EC2 to manage traffic routing and inference.
Why it's wrong here
An external Flask application on EC2 bypasses Databricks Model Serving entirely, so it cannot deliver the endpoint's autoscaling throughput or single-digit-millisecond latency, and MLflow cannot manage the deployment. It tempts when hosting models outside managed infrastructure, but that scenario favours self-managed containers, not Databricks-native serving.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.