Databricks-ML-Pro Model Deployment Practice Question
When deploying a model to a Databricks Model Serving endpoint, what is the purpose of the 'Small', 'Medium', and 'Large' workload size settings?
⚠ Common exam trap
Candidates often mistake these settings for 'number of instances' or 'scaling limits'. They are specifically for compute resource allocation (CPU/RAM) per instance to match model complexity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
They specify the amount of CPU and memory allocated to each model instance.
The workload size setting in Databricks Model Serving defines the compute resources (CPU and Memory) allocated to each instance of the model. Choosing the right size is a trade-off between the complexity of the model's computation and the cost of the infrastructure. Larger models or those with heavy preprocessing requirements need more resources to maintain low latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
They define the maximum number of concurrent requests the endpoint can handle.
Why it's wrong here
The maximum concurrency is determined by the combination of workload size and the number of instances scaled by the autoscaler. The workload size itself refers to the vertical scaling (resource per instance), while the instance count refers to horizontal scaling to handle more concurrent requests.
- ✗
They determine the geographic region where the model will be hosted.
Why it's wrong here
Workload sizes are independent of geographic regions. Region selection is typically handled at the workspace or cloud provider level. The workload size is strictly about the hardware specifications of the virtual machines or containers running the model code within the designated region.
- ✓
They specify the amount of CPU and memory allocated to each model instance.
Why this is correct
Each size tier provides a specific amount of RAM and vCPU. A 'Small' instance might be sufficient for a simple linear regression, while a 'Large' instance would be necessary for complex ensembles or models with large memory footprints to ensure they don't run out of memory during execution.
- ✗
They select the version of the MLflow library used for deployment.
Why it's wrong here
MLflow versions are determined by the environment configuration or the Databricks Runtime version, not by the workload size. The workload size is a resource allocation parameter, whereas the library versions are part of the software environment captured during the model logging process.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.