Courseiva
Model Deployment →hardMultiple Select

Databricks-ML-Pro Model Deployment Practice Question

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure the endpoint can handle sudden spikes in traffic without downtime. The model has a large memory footprint and takes several seconds to load. Which TWO configurations should the engineer implement to achieve this? (Choose two.)

⚠ Common exam trap

The trap here is focusing solely on scaling out quickly while overlooking the need to keep warm instances and provide adequate memory, which are essential for models with slow load times and large footprints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set the minimum provisioned concurrency to a value greater than zero to keep instances warm.

To handle traffic spikes without downtime for a model with a large memory footprint and slow load time, the engineer should keep instances warm by setting a minimum provisioned concurrency greater than zero, and ensure sufficient resources by selecting a larger workload size. These two configurations work together to provide immediate capacity and prevent cold starts, ensuring the endpoint remains responsive under sudden load.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set a low maximum concurrency per instance to force more instances to be created.

    Why it's wrong here

    A low maximum concurrency per instance would limit the number of requests each instance can handle, causing the endpoint to scale out more aggressively. However, this increases the number of instances needed and can lead to higher costs and potential resource contention. For a model with a large memory footprint, having many instances may not be efficient. The goal is to handle spikes, not to artificially increase instance count.

  • ✓

    Set the minimum provisioned concurrency to a value greater than zero to keep instances warm.

    Why this is correct

    Setting a minimum provisioned concurrency ensures that a specified number of model instances are always running, ready to handle requests. This eliminates cold-start latency during traffic spikes, as instances are pre-loaded with the model. For a model with a large memory footprint and slow load time, this is crucial to avoid downtime and maintain performance.

  • ✓

    Configure the endpoint to use a larger workload size to accommodate the model's memory footprint.

    Why this is correct

    A larger workload size provides more memory and compute resources per instance, ensuring the model can be loaded and served efficiently. For a model with a large memory footprint, this prevents out-of-memory errors and allows each instance to handle requests without degradation. Combined with warm instances, this ensures the endpoint can scale to meet traffic spikes.

  • ✗

    Use a smaller model version to reduce load time.

    Why it's wrong here

    Using a smaller model version could reduce load time, but it may compromise model accuracy and is not a configuration change. The scenario assumes the model is fixed; the engineer needs to configure serving parameters. This option does not address the requirement of handling spikes with the given model.

  • ✗

    Enable scale-to-zero to reduce costs during periods of no traffic.

    Why it's wrong here

    Scale-to-zero reduces costs by shutting down all instances when there is no traffic, but it introduces cold starts when traffic resumes. Since the model takes several seconds to load, scaling from zero would cause significant latency and potential downtime during sudden spikes. This configuration is not suitable for the requirement of handling spikes without downtime.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.