Databricks-ML-Assoc Model Deployment Practice Question
A team operates a Databricks Model Serving endpoint with min_instances set to 0 and max_instances set to 4. During a nightly batch job, the endpoint receives a burst of requests and some clients observe elevated latency. The team wants to keep costs low during idle periods while reducing cold-start latency during bursts. Which configuration change best achieves this?
⚠ Common exam trap
Many candidates confuse the roles of min_instances and max_instances, assuming a higher maximum alone removes cold-start latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set min_instances to 1 and leave max_instances at 4 so at least one instance is always warm.
Model Serving endpoints scale between min_instances and max_instances. With min_instances at 0, the endpoint scales to zero when idle and must cold-start a replica when traffic resumes, causing the observed latency. Setting min_instances to 1 keeps a single warm replica available for immediate responses while still allowing scale-out up to max_instances during bursts, achieving both cost and latency goals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set max_instances to 8 so the endpoint can scale further during the burst.
Why it's wrong here
Raising max_instances increases the ceiling for concurrent capacity but does not address the initial cold start when scaling from zero. The first requests still wait for a new instance to provision and load the model, so clients would continue to see elevated latency at the beginning of each burst despite the higher limit.
- ✗
Set min_instances to 4 so the endpoint always matches the maximum capacity.
Why it's wrong here
Setting min_instances equal to max_instances pins four replicas running continuously. This eliminates cold starts but defeats the stated goal of keeping costs low during idle periods, since the endpoint bills for four instances around the clock even when no traffic arrives. It over-provisions for a nightly burst.
- ✓
Set min_instances to 1 and leave max_instances at 4 so at least one instance is always warm.
Why this is correct
Keeping min_instances at 1 preserves a warm replica that can answer requests immediately, eliminating scale-from-zero cold starts while max_instances still caps cost during spikes. This balances the cost-saving goal during idle periods with the latency goal during bursts, because only one instance runs continuously rather than four.
- ✗
Enable scale-to-zero by setting both min_instances and max_instances to 0 and rely on queueing.
Why it's wrong here
Setting max_instances to 0 would leave the endpoint with no capacity to serve requests at all. Scale-to-zero applies to the minimum, not the maximum; the maximum must be at least one for the endpoint to function. This configuration would cause requests to fail rather than queue, breaking serving entirely.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.