Databricks-GenAI-Assoc Application Development Practice Question
Exhibit
{
"model_name": "my_model",
"model_version": "1",
"config": {
"served_models": [
{
"model_name": "my_model",
"model_version": "1",
"workload_type": "CPU",
"scale_to_zero_enabled": true
}
]
}
}Refer to the exhibit. A developer wants to update this serving endpoint configuration to ensure it handles high-concurrency requests with consistent latency. Which change should be applied to the configuration?
⚠ Common exam trap
Candidates frequently select scale-to-zero options when high concurrency and consistent low latency are required, confusing cost savings with strict performance SLO requirements.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set 'scale_to_zero_enabled' to 'false' and define 'min_provisioned_concurrency'.
To handle high-concurrency with consistent latency, the configuration must move away from 'scale_to_zero_enabled: true' and utilize 'min_provisioned_concurrency'. Scaling to zero introduces cold-start latency that disrupts consistent response times required for high-concurrency production applications. By pinning minimum resources, the model remains active and ready to handle incoming traffic immediately, ensuring that performance remains predictable even during intermittent bursts of user activity within the production environment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Change 'workload_type' to 'GPU' and remove 'scale_to_zero_enabled'.
Why it's wrong here
GPU acceleration is only beneficial for specific model types like LLMs or deep learning architectures. Switching to GPU without verifying model requirements incurs unnecessary costs and does not address the latency concerns related to scaling behavior. Workload type must align with the specific model's compute needs for efficiency.
- ✓
Set 'scale_to_zero_enabled' to 'false' and define 'min_provisioned_concurrency'.
Why this is correct
Disabling scale-to-zero prevents the endpoint from shutting down, which eliminates cold-start latency. By defining 'min_provisioned_concurrency', the developer reserves a set number of replicas that are always running, ensuring that the system can handle concurrent requests immediately without waiting for infrastructure initialization, thus stabilizing latency under heavy load.
- ✗
Increase the 'model_version' to 'latest' to trigger automatic load balancing.
Why it's wrong here
Changing the model version to 'latest' does not manage concurrency or latency; it merely points the endpoint to a different model artifact. Versioning is a lifecycle management strategy, not an infrastructure scaling strategy. It does not resolve issues related to cold starts or the need for consistent throughput.
- ✗
Add a 'timeout' parameter to the JSON configuration block.
Why it's wrong here
Adding a timeout parameter dictates the maximum time a request can take before failing, but it does not improve performance or latency. In fact, it might increase the rate of request failures under high concurrency without actually optimizing the underlying compute capacity required for smoother execution.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.