Courseiva
Model Deployment →easyMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

A data scientist has deployed a model to a Databricks Model Serving endpoint. The endpoint is configured with scale-to-zero enabled and a workload size of Small. After a period of inactivity, the endpoint scales down to zero. A client application sends a request to the endpoint after this idle period. What happens to the first request?

⚠ Common exam trap

The trap here is assuming that scale-to-zero causes request failures or that a warm fallback exists, when in fact the request is queued and served after a cold start.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The request is queued until the endpoint scales up, then processed with increased latency.

With scale-to-zero, the endpoint reduces replicas to zero after inactivity. When a new request arrives, the system must provision resources and load the model, causing the request to be queued and served with higher latency. This is expected behavior and not an error.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The request fails with a 503 Service Unavailable error because the endpoint is offline.

    Why it's wrong here

    Scale-to-zero does not make the endpoint permanently unavailable; it simply reduces replicas to zero. The service automatically scales up when a request arrives. The request is not rejected outright; it is queued. A 503 error would occur only if the endpoint is deleted or misconfigured.

  • ✗

    The request is immediately served by a cold-start replica with no added latency.

    Why it's wrong here

    Scale-to-zero means no replicas are running, so there is no warm replica to serve the request instantly. A cold start is required, which introduces latency. The claim of no added latency is incorrect; cold starts inherently add delay while the model loads and resources provision.

  • ✓

    The request is queued until the endpoint scales up, then processed with increased latency.

    Why this is correct

    When scale-to-zero is enabled, the endpoint scales down to zero replicas after inactivity. The first request after idle time triggers a scale-up, causing the request to be queued until a replica is available. This results in higher latency for that initial request. Subsequent requests are served with normal latency.

  • ✗

    The request is routed to a fallback model version that is always kept warm.

    Why it's wrong here

    Databricks Model Serving does not automatically maintain a fallback warm model version. Scale-to-zero applies to the entire endpoint, and all served entities scale down. There is no built-in fallback mechanism that keeps a model warm. The request will be queued until the primary model scales up.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.