Courseiva
ML Ops →mediumMultiple Choice

Databricks-ML-Pro ML Ops Practice Question

When deploying a model to a Databricks Model Serving endpoint, how can you ensure the model scales automatically based on traffic demand?

⚠ Common exam trap

Candidates often confuse manual endpoint configuration with auto-scaling, mistakenly believing they must manually add nodes or adjust cluster sizes during traffic spikes in the Databricks UI.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure auto-scaling settings in the Model Serving endpoint definition to handle traffic fluctuations.

Databricks Model Serving utilizes auto-scaling compute resources to handle fluctuations in traffic. By configuring the serving endpoint with appropriate auto-scaling parameters, the platform dynamically adjusts the number of concurrent instances based on request load. This ensures that the application maintains low latency under heavy load while remaining cost-effective during periods of low activity, which is essential for managing production-level model inference at scale.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Manually adjust the number of instances in the endpoint configuration every time traffic increases.

    Why it's wrong here

    Manual scaling is reactive and impractical in production environments. It leads to service latency when traffic spikes unexpectedly, as there is a delay before the operator can intervene. Scaling should be handled by the platform's native auto-scaling capabilities to ensure responsiveness and operational efficiency.

  • ✗

    Set the endpoint to a fixed number of instances that is high enough to handle peak traffic.

    Why it's wrong here

    Setting a fixed high instance count is a waste of resources and significantly increases costs during off-peak hours. It does not adapt to changing demand, and if traffic exceeds the pre-provisioned capacity, the system will still face latency issues or request failures, defeating the purpose of efficient infrastructure.

  • ✓

    Configure auto-scaling settings in the Model Serving endpoint definition to handle traffic fluctuations.

    Why this is correct

    Auto-scaling is a core feature of Databricks Model Serving. By enabling this configuration, the platform automatically manages the allocation of compute resources to match incoming traffic patterns. This provides a balance between high availability during peaks and cost optimization during periods of lower utilization without manual intervention.

  • ✗

    Use a load balancer to redirect traffic to different model endpoints depending on the time of day.

    Why it's wrong here

    Routing traffic based on time is an unreliable proxy for actual demand. It does not account for sudden spikes, unusual traffic patterns, or model availability. It adds unnecessary complexity to the infrastructure, whereas native auto-scaling is designed to handle this logic more efficiently and accurately.

About these practice questions

One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.