Databricks-ML-Assoc Model Deployment Practice Question
A team is preparing to deploy a model to a Databricks Model Serving endpoint that must scale down to zero replicas when idle yet still serve bursty traffic with acceptable cold-start latency. They also need to capture the request and response payloads for later monitoring. Which two endpoint settings or features should they configure? (Choose two.)
⚠ Common exam trap
The trap here is assuming that reducing cold-start latency requires keeping a warm replica, when that choice directly prevents the endpoint from ever scaling down to zero as required.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable scale-to-zero on the served entity so the endpoint can drop to zero replicas during idle periods.
The scenario has two distinct requirements: scaling to zero when idle and capturing payloads for monitoring. Scale-to-zero satisfies the first by releasing all compute during idle periods, and inference tables satisfy the second by logging request and response payloads to a Delta table. Keeping a minimum replica or reserving provisioned throughput would prevent scaling to zero, and workload size does not address either requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure a provisioned throughput endpoint to reserve dedicated capacity for the model.
Why it's wrong here
Provisioned throughput reserves dedicated capacity to deliver predictable, low-latency performance, but it keeps resources allocated and is not compatible with scaling down to zero replicas when idle. The scenario prioritizes scaling to zero, so reserving dedicated capacity conflicts with that goal. Provisioned throughput is better suited to steady, latency-sensitive workloads that must avoid cold starts entirely.
- ✗
Increase the workload size to Large to reduce cold-start latency when scaling up from zero.
Why it's wrong here
Workload size controls the CPU and memory allocated per replica, not how quickly the endpoint initializes after scaling from zero. A larger workload size may even lengthen container startup because more resources must be provisioned. It does not provide payload logging, and it does not by itself enable scaling to zero. The scenario's two requirements are met by scale-to-zero and inference tables, not by resizing the workload.
- ✓
Enable scale-to-zero on the served entity so the endpoint can drop to zero replicas during idle periods.
Why this is correct
Scale-to-zero lets the endpoint release all compute when no requests arrive, which directly satisfies the requirement to scale down to zero replicas when idle. When traffic resumes, the endpoint scales back up, incurring a cold start. Because the scenario demands scaling to zero, this setting is necessary; without it the endpoint would keep at least one replica running and continue consuming capacity during idle periods.
- ✓
Enable inference tables on the endpoint to log request and response payloads to a Delta table.
Why this is correct
Inference tables automatically capture the request payloads and the model's responses and write them to a Delta table for monitoring and analysis. This directly fulfills the requirement to record payloads for later inspection. Enabling the feature requires specifying a destination table, after which the endpoint logs each scored request, making it the correct choice for the observability portion of the scenario.
- ✗
Set the endpoint's minimum replicas to one so that a warm replica is always available.
Why it's wrong here
Keeping a minimum of one replica prevents the endpoint from ever reaching zero, which contradicts the explicit requirement to scale down to zero when idle. While it would reduce cold-start latency by keeping compute warm, it sacrifices the cost-saving behavior the team asked for. This setting is the opposite of scale-to-zero and therefore cannot be one of the correct choices here.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.