Databricks-GenAI-Assoc Assembling and Deploying Apps Practice Question
Exhibit
{"model": "llama-3", "config": {"gpu": "nvidia_a10", "min_instances": 0}, "autoscaling": "enabled"}Refer to the exhibit. What is the impact of min_instances: 0 on this deployment?
⚠ Common exam trap
Candidates often fear that 'min_instances: 0' will cause the model to be deleted or unavailable, failing to realize it is a standard cost-saving feature for serverless endpoints.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The endpoint will scale to zero when idle, reducing costs.
Setting min_instances to 0 in Databricks Model Serving enables 'scale-to-zero'. This feature is highly effective for cost optimization, as it automatically shuts down the GPU resources when no traffic is present. Upon receiving a new request, the endpoint will automatically scale up, though this may introduce a slight 'cold start' latency. This trade-off is often acceptable for non-critical or low-traffic services to drastically reduce infrastructure costs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The endpoint will be permanently disabled.
Why it's wrong here
The endpoint is not disabled; it is simply scaled down to zero resources when idle. It remains active and ready to accept traffic, but the infrastructure resources are only provisioned on demand. This is a standard and well-supported feature for optimizing the financial efficiency of Databricks Model Serving.
- ✓
The endpoint will scale to zero when idle, reducing costs.
Why this is correct
This configuration allows the endpoint to release all compute resources when no requests are being processed. This is a best practice for cost management in AI deployments, as it prevents paying for expensive GPU instances when they are not actively serving traffic, enabling more sustainable and efficient cloud usage.
- ✗
The endpoint will have higher latency than a fixed instance.
Why it's wrong here
Cold start latency only affects the first request after the endpoint has scaled down. It does not mean the overall latency of the model is higher during normal operation. Once instances are running, performance is identical to a fixed-size deployment, so this is not a general performance increase.
- ✗
The endpoint will be restricted to CPU usage only.
Why it's wrong here
The choice of GPU instance type is independent of the auto-scaling policy. Setting min_instances to zero does not change the underlying hardware or the model's ability to utilize GPU acceleration once it is provisioned. The configuration merely controls the scaling behavior, not the hardware capabilities of the endpoint.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.