Databricks-GenAI-Assoc Assembling and Deploying Apps Practice Question
A team deploys a Mosaic AI Agent application to a Databricks Model Serving endpoint. During load testing they observe that the first request after an idle period takes several seconds, while subsequent requests are fast. They want to eliminate this cold-start penalty for a latency-sensitive customer-facing application while keeping costs reasonable during off-peak hours. Which configuration should they apply?
⚠ Common exam trap
Many exam-takers confuse concurrency limits or observability features with warm capacity, when only provisioned concurrency with a nonzero minimum prevents the endpoint from scaling fully to zero.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the endpoint's minimum provisioned concurrency to a value greater than zero so at least one instance is always warm.
Cold starts occur when the endpoint has scaled down and must load a fresh model container on the next request. Provisioned concurrency with a minimum above zero keeps at least one instance warm at all times, removing the first-request latency penalty while still permitting scale-up under load and scale-down to the configured minimum during off-peak hours. Concurrency, inference tables, and image size do not guarantee warm capacity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set the endpoint's minimum provisioned concurrency to a value greater than zero so at least one instance is always warm.
Why this is correct
Provisioned concurrency keeps a specified number of model instances loaded and ready, so requests never wait for a cold container to load the agent and its dependencies. Setting the minimum above zero removes the idle-time cold-start penalty while still allowing the endpoint to scale up under load. This is the intended control for latency-sensitive endpoints that must stay responsive.
- ✗
Reduce the agent's dependency footprint and re-log the model so the container image is smaller.
Why it's wrong here
A smaller image can shorten load time somewhat, but the endpoint can still scale to zero during idle periods, so the first request after idle still pays a load cost. This is a partial mitigation, not an elimination of the cold-start penalty. It also requires re-validating the agent, and it does not give the deterministic warm capacity the scenario demands.
- ✗
Enable inference tables on the endpoint to log requests and responses, which keeps the model warm.
Why it's wrong here
Inference tables capture payloads for monitoring and debugging; they do not keep model instances loaded or influence scaling decisions. Enabling them adds storage and governance overhead without changing cold-start behavior. This option conflates observability with capacity management, which are separate concerns on Model Serving.
- ✗
Increase the endpoint's maximum concurrency per instance so a single warm instance can absorb all traffic.
Why it's wrong here
Raising max concurrency changes how many simultaneous requests one instance handles, not whether an instance is already loaded. After an idle period the endpoint may still have scaled to zero and must load a fresh container, so the first-request penalty remains. Concurrency tuning helps throughput under sustained load but does not address cold starts.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.