Courseiva
Model Deployment →hardMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

A team has deployed a model to Databricks Model Serving and wants to enable autoscaling to handle variable traffic. They configure the endpoint with scale_to_zero_enabled set to true and a min_provisioned_concurrency of 0. After deployment, they notice that the endpoint takes several seconds to respond to the first request after a period of inactivity. What is the cause of this latency?

⚠ Common exam trap

The trap here is attributing cold-start latency to model artifact retrieval or environment setup instead of the scale-from-zero provisioning process.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The endpoint is scaling from zero, which requires provisioning resources and loading the model before serving the request.

Enabling scale_to_zero with zero minimum concurrency allows the endpoint to shut down completely when idle, reducing cost. However, the next request must wait for compute resources to be provisioned and the model to be loaded, resulting in cold-start latency. To avoid this, set a min_provisioned_concurrency greater than zero or disable scale_to_zero, keeping at least one replica warm.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model is being reloaded from Unity Catalog on each request, causing cold-start latency.

    Why it's wrong here

    Model artifacts are cached after the initial load; they are not reloaded from Unity Catalog on every request. The latency described is due to scaling from zero, not artifact retrieval. Unity Catalog access is only part of the initial provisioning, not the per-request path, so this explanation does not match the observed behavior.

  • ✗

    The endpoint is using a GPU workload, and GPU initialization takes several seconds after each idle period.

    Why it's wrong here

    GPU initialization can add latency during cold starts, but the scenario does not specify a GPU workload. The primary cause of latency after inactivity with scale_to_zero_enabled is the scaling from zero, regardless of CPU or GPU. GPU initialization is a secondary factor and not the root cause described here, making this option less accurate.

  • ✗

    The model's Python environment is being reinstalled on each request due to missing dependencies in the container image.

    Why it's wrong here

    The Python environment is built into the model's container image and is not reinstalled per request. Dependencies are resolved during endpoint creation or update, not at inference time. While environment issues can cause deployment failures, they do not produce intermittent latency after idle periods. The described symptom aligns with scaling from zero, not environment reinstallation.

  • ✓

    The endpoint is scaling from zero, which requires provisioning resources and loading the model before serving the request.

    Why this is correct

    When scale_to_zero_enabled is true and min_provisioned_concurrency is 0, the endpoint scales down to zero replicas during inactivity. The first request after idle time triggers a cold start: the system must provision compute resources and load the model into memory. This provisioning and loading time causes the several-second latency observed, which is inherent to scale-to-zero behavior.

About these practice questions

This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.