Databricks-ML-Pro Model Deployment Practice Question
Exhibit
GET /serving-endpoints/customer-churn/config
{
"name": "customer-churn",
"config": {
"served_entities": [
{
"entity_name": "prod.models.churn",
"entity_version": "5",
"scale_to_zero_enabled": false
}
]
}
}Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?
⚠ Common exam trap
Candidates often assume endpoints automatically scale down to zero when idle, forgetting that scale-to-zero must be explicitly enabled to avoid continuous compute charges.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.
The configuration shows that 'scale_to_zero_enabled' is set to false. This means that at least one instance of the model is running 24/7, regardless of whether any requests are being made. While this ensures there is never a cold start, it leads to continuous billing for the compute resources even during nights and weekends.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The 'entity_version' is set to a legacy version that uses more expensive hardware.
Why it's wrong here
Hardware is determined by the workload size (which is not shown but defaults to Small), not by the model version number. Version 5 and Version 1 would run on the same underlying instance type unless the workload_size parameter itself was modified between deployments.
- ✓
The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.
Why this is correct
When scale-to-zero is disabled, the system maintains the minimum number of provisioned instances (defaulting to 1). This provides the benefit of zero latency for the first request after an idle period but results in constant resource consumption and associated cloud costs.
- ✗
The endpoint is using a GPU-accelerated instance by default for all Unity Catalog models.
Why it's wrong here
Databricks Model Serving does not default to GPU instances; users must explicitly select a GPU workload type. The exhibit does not show a GPU type being requested, so the cost is likely coming from the continuous running of standard CPU-based instances instead.
- ✗
Multiple versions of the model are being served simultaneously, doubling the cost.
Why it's wrong here
The 'served_entities' list in the exhibit contains only one object (entity_version 5). Therefore, only one model version is active. If multiple versions were being served, they would each appear as separate entries in the 'served_entities' array within the JSON configuration.
About these practice questions
One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.