Courseiva
ML Ops →mediumMultiple Choice

Databricks-ML-Pro ML Ops Practice Question

When deploying a model as a real-time REST endpoint on Databricks, how can you ensure the infrastructure scales automatically to handle increased request traffic?

⚠ Common exam trap

Candidates often choose manual cluster configuration or custom scaling scripts, failing to realize that Model Serving endpoints provide native, serverless auto-scaling without needing manual infrastructure overhead.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy the model using Databricks Model Serving endpoints.

Databricks Model Serving provides serverless, auto-scaling inference endpoints. By default, it manages the underlying infrastructure, scaling the number of replicas based on real-time request volume. This is a key MLOps requirement to ensure high availability and performance without manual intervention, allowing teams to handle unpredictable traffic spikes without over-provisioning and incurring unnecessary costs during low-traffic periods.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure the cluster to use a fixed number of workers in the cluster settings.

    Why it's wrong here

    Fixed-size clusters do not auto-scale based on request traffic. This leads to either under-provisioning during peak times, causing request timeouts, or over-provisioning during off-peak times, resulting in wasted compute costs. Auto-scaling is a necessary feature for production-ready, cost-efficient, and resilient machine learning model deployment on Databricks.

  • ✓

    Deploy the model using Databricks Model Serving endpoints.

    Why this is correct

    Databricks Model Serving endpoints are designed specifically for high-performance, real-time inference. They automatically handle infrastructure provisioning and scaling, adjusting the number of active model instances based on the current load. This abstracts away the complexity of cluster management, ensuring reliable and efficient model serving in production.

  • ✗

    Manually add instances to the inference cluster using the Databricks API before peak times.

    Why it's wrong here

    Manual scaling is reactive and prone to human error. It fails to address sudden traffic spikes and requires constant monitoring. Modern MLOps practices emphasize automated, elastic infrastructure that responds to real-time metrics, rather than manual intervention, which is neither scalable nor sustainable in an enterprise data environment.

  • ✗

    Enable 'Auto-terminate' on the inference cluster configuration.

    Why it's wrong here

    Auto-terminate is a cost-saving feature that shuts down idle clusters. It does not provide the ability to scale infrastructure based on request volume. In fact, if auto-terminate is enabled on a serving cluster, it might shut down the service entirely, leading to downtime and service unavailability for users.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.