Databricks-ML-Pro ML Ops Practice Question
When deploying a model as a real-time REST endpoint on Databricks, how can you ensure the infrastructure scales automatically to handle increased request traffic?
⚠ Common exam trap
Candidates often choose manual cluster configuration or custom scaling scripts, failing to realize that Model Serving endpoints provide native, serverless auto-scaling without needing manual infrastructure overhead.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the model using Databricks Model Serving endpoints.
Databricks Model Serving provides serverless, auto-scaling inference endpoints. By default, it manages the underlying infrastructure, scaling the number of replicas based on real-time request volume. This is a key MLOps requirement to ensure high availability and performance without manual intervention, allowing teams to handle unpredictable traffic spikes without over-provisioning and incurring unnecessary costs during low-traffic periods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure the cluster to use a fixed number of workers in the cluster settings.
Why it's wrong here
Fixed-size clusters do not auto-scale based on request traffic. This leads to either under-provisioning during peak times, causing request timeouts, or over-provisioning during off-peak times, resulting in wasted compute costs. Auto-scaling is a necessary feature for production-ready, cost-efficient, and resilient machine learning model deployment on Databricks.
- ✓
Deploy the model using Databricks Model Serving endpoints.
Why this is correct
Databricks Model Serving endpoints are designed specifically for high-performance, real-time inference. They automatically handle infrastructure provisioning and scaling, adjusting the number of active model instances based on the current load. This abstracts away the complexity of cluster management, ensuring reliable and efficient model serving in production.
- ✗
Manually add instances to the inference cluster using the Databricks API before peak times.
Why it's wrong here
Manual scaling is reactive and prone to human error. It fails to address sudden traffic spikes and requires constant monitoring. Modern MLOps practices emphasize automated, elastic infrastructure that responds to real-time metrics, rather than manual intervention, which is neither scalable nor sustainable in an enterprise data environment.
- ✗
Enable 'Auto-terminate' on the inference cluster configuration.
Why it's wrong here
Auto-terminate is a cost-saving feature that shuts down idle clusters. It does not provide the ability to scale infrastructure based on request volume. In fact, if auto-terminate is enabled on a serving cluster, it might shut down the service entirely, leading to downtime and service unavailability for users.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.