Courseiva
Model Deployment →hardMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

Exhibit

GET /serving-endpoints/customer-churn/config
{
  "name": "customer-churn",
  "config": {
    "served_entities": [
      {
        "entity_name": "prod.models.churn",
        "entity_version": "5",
        "scale_to_zero_enabled": false
      }
    ]
  }
}

Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?

⚠ Common exam trap

Candidates often assume endpoints automatically scale down to zero when idle, forgetting that scale-to-zero must be explicitly enabled to avoid continuous compute charges.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.

The configuration shows that 'scale_to_zero_enabled' is set to false. This means that at least one instance of the model is running 24/7, regardless of whether any requests are being made. While this ensures there is never a cold start, it leads to continuous billing for the compute resources even during nights and weekends.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The 'entity_version' is set to a legacy version that uses more expensive hardware.

    Why it's wrong here

    Hardware is determined by the workload size (which is not shown but defaults to Small), not by the model version number. Version 5 and Version 1 would run on the same underlying instance type unless the workload_size parameter itself was modified between deployments.

  • ✓

    The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.

    Why this is correct

    When scale-to-zero is disabled, the system maintains the minimum number of provisioned instances (defaulting to 1). This provides the benefit of zero latency for the first request after an idle period but results in constant resource consumption and associated cloud costs.

  • ✗

    The endpoint is using a GPU-accelerated instance by default for all Unity Catalog models.

    Why it's wrong here

    Databricks Model Serving does not default to GPU instances; users must explicitly select a GPU workload type. The exhibit does not show a GPU type being requested, so the cost is likely coming from the continuous running of standard CPU-based instances instead.

  • ✗

    Multiple versions of the model are being served simultaneously, doubling the cost.

    Why it's wrong here

    The 'served_entities' list in the exhibit contains only one object (entity_version 5). Therefore, only one model version is active. If multiple versions were being served, they would each appear as separate entries in the 'served_entities' array within the JSON configuration.

About these practice questions

One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.