Courseiva

Databricks-GenAI-Assoc Assembling and Deploying Apps Practice Question

Exhibit

{"model": "llama-3", "config": {"gpu": "nvidia_a10", "min_instances": 0}, "autoscaling": "enabled"}

Refer to the exhibit. What is the impact of min_instances: 0 on this deployment?

⚠ Common exam trap

Candidates often fear that 'min_instances: 0' will cause the model to be deleted or unavailable, failing to realize it is a standard cost-saving feature for serverless endpoints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The endpoint will scale to zero when idle, reducing costs.

Setting min_instances to 0 in Databricks Model Serving enables 'scale-to-zero'. This feature is highly effective for cost optimization, as it automatically shuts down the GPU resources when no traffic is present. Upon receiving a new request, the endpoint will automatically scale up, though this may introduce a slight 'cold start' latency. This trade-off is often acceptable for non-critical or low-traffic services to drastically reduce infrastructure costs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The endpoint will be permanently disabled.

    Why it's wrong here

    The endpoint is not disabled; it is simply scaled down to zero resources when idle. It remains active and ready to accept traffic, but the infrastructure resources are only provisioned on demand. This is a standard and well-supported feature for optimizing the financial efficiency of Databricks Model Serving.

  • ✓

    The endpoint will scale to zero when idle, reducing costs.

    Why this is correct

    This configuration allows the endpoint to release all compute resources when no requests are being processed. This is a best practice for cost management in AI deployments, as it prevents paying for expensive GPU instances when they are not actively serving traffic, enabling more sustainable and efficient cloud usage.

  • ✗

    The endpoint will have higher latency than a fixed instance.

    Why it's wrong here

    Cold start latency only affects the first request after the endpoint has scaled down. It does not mean the overall latency of the model is higher during normal operation. Once instances are running, performance is identical to a fixed-size deployment, so this is not a general performance increase.

  • ✗

    The endpoint will be restricted to CPU usage only.

    Why it's wrong here

    The choice of GPU instance type is independent of the auto-scaling policy. Setting min_instances to zero does not change the underlying hardware or the model's ability to utilize GPU acceleration once it is provisioned. The configuration merely controls the scaling behavior, not the hardware capabilities of the endpoint.

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.