Databricks-ML-Pro Model Deployment Practice Question
A data science team is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. They want to configure auto-scaling appropriately. Which TWO parameters should they adjust to control the scaling behavior? (Choose two.)
⚠ Common exam trap
The trap here is focusing on cost-saving features like scale-to-zero, which actually hinder burst handling, instead of the core scaling parameters.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the minimum number of replicas to a value greater than zero.
Auto-scaling in Databricks Model Serving is governed by the minimum and maximum replica counts. The minimum ensures baseline capacity to handle initial bursts, while the maximum allows the endpoint to scale out to meet high demand. These two parameters directly control the scaling range and are essential for handling traffic spikes without dropping requests.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable scale-to-zero to reduce costs during idle periods.
Why it's wrong here
Scale-to-zero reduces cost by scaling down to zero when idle, but it introduces cold start latency when traffic resumes. For handling sudden spikes without dropping requests, scale-to-zero is counterproductive because it may cause initial requests to queue or time out during scale-up.
- ✗
Adjust the model's batch size in the predict function to process more requests per inference.
Why it's wrong here
Batch size affects the model's inference throughput but does not control the endpoint's auto-scaling behavior. While optimizing batch size can improve efficiency, it does not directly influence how many replicas are provisioned in response to traffic.
- ✓
Set the minimum number of replicas to a value greater than zero.
Why this is correct
Setting a minimum replica count ensures that a baseline capacity is always available, reducing cold start impact and allowing the endpoint to handle initial bursts without waiting for scaling. This is crucial for latency-sensitive applications that cannot tolerate cold starts.
- ✓
Configure the maximum number of replicas to a high value to allow scaling out.
Why this is correct
The maximum replica count defines the upper limit of scaling. Setting it high allows the endpoint to scale out to accommodate large traffic spikes, preventing request drops due to insufficient capacity. However, it should be balanced with cost considerations.
- ✗
Set the endpoint's workload size to 'Large' to increase per-replica capacity.
Why it's wrong here
Workload size determines the compute resources (CPU, memory) allocated per replica, which can affect the number of concurrent requests a single replica can handle. However, it does not directly control the scaling parameters like minimum and maximum replicas, which are the primary levers for auto-scaling.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.