Courseiva
Model Deployment →hardMultiple Choice

Databricks-ML-Assoc Model Deployment Practice Question

A team has a Databricks Model Serving endpoint configured with scale-to-zero enabled and min_instances set to 0. During a load test, they observe that the first request after an idle period takes roughly 40 seconds while subsequent requests complete in under 200 milliseconds. They need to eliminate this cold-start latency for a customer-facing application without over-provisioning. Which configuration change best addresses the requirement?

⚠ Common exam trap

Many candidates confuse max_instances, which governs burst capacity, with min_instances, which governs whether any replica stays warm when traffic drops to zero.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set min_instances to 1 so at least one instance stays warm.

Cold-start latency occurs because a scaled-to-zero endpoint must provision a container and load the model before serving the first request. Keeping one replica always available by setting min_instances to 1 removes that provisioning and loading delay while still allowing the endpoint to scale up under load. Increasing max_instances, changing workload size, or enabling routing features does not prevent the endpoint from scaling to zero during idle periods.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase max_instances so more replicas can absorb the burst.

    Why it's wrong here

    max_instances controls how many replicas can be added under load; it does not keep any replica warm when traffic is zero. With scale-to-zero and min_instances at 0, the endpoint still scales down to nothing, so raising max_instances only helps once a cold start has already occurred and traffic ramps. The 40-second first-request latency would remain unchanged for the initial request.

  • ✗

    Increase the workload size from Small to Large.

    Why it's wrong here

    Workload size determines the CPU and memory allocated per replica, which affects throughput and memory headroom for large models, not whether a replica stays warm. A Large workload still scales to zero when min_instances is 0, so the 40-second cold start persists. Choosing a larger workload would only increase cost and may improve per-request performance once warm, but it does not solve the idle-time latency problem.

  • ✗

    Enable route optimization in the endpoint configuration.

    Why it's wrong here

    Route optimization reduces overhead in routing requests to replicas but does not pre-provision compute or keep a model loaded. Cold-start latency comes from provisioning a container and loading model artifacts, which route optimization cannot eliminate. There is no Databricks Model Serving setting called route optimization that addresses idle scale-down, so this change would not fix the observed first-request delay.

  • ✓

    Set min_instances to 1 so at least one instance stays warm.

    Why this is correct

    Setting min_instances to 1 keeps one replica provisioned at all times, so the container and model are already loaded and the first request after an idle period is served without a cold start. This is the documented way to trade a small amount of always-on compute for predictable low latency, and it does not require scaling the endpoint beyond what the workload needs during quiet periods.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.