Courseiva
Model Deployment →hardMultiple Select

Databricks-ML-Assoc Model Deployment Practice Question

A team owns a Databricks Model Serving endpoint that receives sporadic bursts of traffic. They want to reduce cold-start latency during bursts while keeping cost predictable, and they also need to capture the request payloads and predictions for later monitoring. Which two configuration choices should they make? (Choose two.)

⚠ Common exam trap

The trap here is treating scale_to_zero_enabled, which saves cost by idling replicas, as a latency optimization, when it is actually the primary cause of cold starts on a serving endpoint.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set min_instances to a value greater than zero so a warm replica is always available.

Cold-start latency during sporadic bursts is best mitigated by keeping at least one replica warm with min_instances greater than zero, avoiding scale-to-zero teardown. Capturing request payloads and predictions for monitoring is done with inference tables, which write each request and response to a Delta table. Together these meet both the latency and the observability goals without conflicting.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Set scale_to_zero_enabled to true on the served entity to lower idle cost.

    Why it's wrong here

    Scale to zero reduces cost by shutting down replicas when there is no traffic, but it is precisely what causes cold starts: the first request after an idle period must wait for a replica to spin up and load the model. Enabling it works against the stated goal of reducing burst latency, so it is the wrong choice here.

  • ✗

    Increase the endpoint's workload size from Small to Large.

    Why it's wrong here

    Workload size controls the compute resources, such as CPU and memory, allocated to each replica and thus affects throughput and per-request latency under load. It does not eliminate the delay of starting a replica from zero, and it does not log any data. Choosing a larger size can raise cost without solving the cold-start or logging requirements.

  • ✓

    Set min_instances to a value greater than zero so a warm replica is always available.

    Why this is correct

    Setting min_instances above zero keeps at least that many replicas provisioned and ready, so requests arriving after an idle period do not wait for a new container to load the model. This directly addresses cold-start latency during bursts, at the cost of paying for the always-on capacity. It does not by itself record request payloads or predictions.

  • ✗

    Attach an MLflow experiment to the endpoint to record prediction requests.

    Why it's wrong here

    MLflow experiments track training runs and their metrics, parameters, and artifacts; they are not a serving-time logging mechanism. You cannot attach an experiment to a Model Serving endpoint to capture live request payloads. Inference tables are the Databricks feature that records served requests and responses for monitoring.

  • ✓

    Enable inference tables on the endpoint so requests and responses are logged to a Delta table.

    Why this is correct

    Inference tables automatically capture the request payload, the response, and metadata such as timestamp and status for each scored request, writing them to a Delta table in Unity Catalog. That satisfies the requirement to retain payloads and predictions for monitoring and later analysis. It has no effect on cold-start behavior, so it complements rather than replaces warm capacity.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.