Courseiva
Model Deployment →hardMultiple Select

Databricks-ML-Pro Model Deployment Practice Question

An ML engineer is deploying a model to Databricks Model Serving and needs to enable automatic scaling based on traffic. The model has variable inference latency and the team wants to optimize cost while maintaining performance. Which TWO configurations are required to achieve this? (Choose two.)

⚠ Common exam trap

Many candidates confuse cost-saving features like scale-to-zero with autoscaling, which requires replica bounds and a concurrency target.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable auto-scaling by specifying a target concurrency per replica.

To enable automatic scaling for a Databricks Model Serving endpoint, you must define the minimum and maximum replica counts and set a target concurrency per replica. These settings allow the serving infrastructure to adjust the number of replicas in response to traffic, ensuring performance while controlling cost. Other options either do not directly enable scaling or are not applicable to Model Serving.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model using a GPU-enabled workload type to reduce latency.

    Why it's wrong here

    GPU acceleration can improve inference speed but does not enable automatic scaling. The question asks for configurations to achieve autoscaling, not to reduce latency. Using a GPU workload type is orthogonal to scaling and may increase cost without addressing the scaling requirement. It is not a required configuration for autoscaling.

  • ✓

    Enable auto-scaling by specifying a target concurrency per replica.

    Why this is correct

    Auto-scaling in Databricks Model Serving is driven by a target concurrency metric. You specify the desired number of concurrent requests per replica, and the system adjusts the replica count to maintain that target. This is essential for scaling based on traffic, as it directly ties replica count to request load.

  • ✓

    Configure the minimum and maximum number of replicas for the endpoint.

    Why this is correct

    Databricks Model Serving allows you to set a minimum and maximum replica count to bound the autoscaling behavior. The minimum ensures baseline capacity, while the maximum prevents over-provisioning. This is a core requirement for enabling automatic scaling based on traffic, as the system scales between these bounds according to load.

  • ✗

    Set the endpoint to use a dedicated cluster with autoscaling enabled.

    Why it's wrong here

    Databricks Model Serving abstracts the underlying compute and manages scaling automatically; you do not configure a dedicated cluster with autoscaling. The serving infrastructure handles resource allocation based on the replica settings. Specifying a dedicated cluster is not a supported configuration for Model Serving endpoints.

  • ✗

    Set the scale-to-zero option to true so the endpoint can shut down when idle.

    Why it's wrong here

    Scale-to-zero is not supported for all workload types and can introduce cold-start latency. While it reduces cost during idle periods, it does not directly address scaling based on traffic fluctuations. The question focuses on automatic scaling under load, not idle shutdown. Enabling scale-to-zero alone would not provide the necessary scaling behavior for variable traffic.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.