Databricks-GenAI-Assoc Design Applications Practice Question
You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?
⚠ Common exam trap
Candidates often select standard auto-scaling for high-concurrency needs, failing to realize that Provisioned Throughput is specifically required to guarantee performance for high-load production LLM workloads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable 'Provisioned Throughput' on the serving endpoint configuration.
Enabling Provisioned Throughput for Model Serving allows for dedicated resources that can handle high-concurrency workloads efficiently. This is crucial for applications serving many users simultaneously. By optimizing resource allocation, you prevent bottlenecks during peak usage, ensuring consistent performance and minimizing request latency, which directly impacts the user experience and the scalability of the overall application architecture.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the serving endpoint to use CPU-only instances for all LLM inference tasks.
Why it's wrong here
CPU-only instances are generally unsuitable for high-concurrency LLM inference because LLMs are computationally intensive. Using CPUs will result in significant latency and low throughput, failing to meet the requirements for a responsive, high-scale application that needs to handle many simultaneous requests effectively.
- ✓
Enable 'Provisioned Throughput' on the serving endpoint configuration.
Why this is correct
Provisioned Throughput provides dedicated capacity for foundation models, which is essential for high-concurrency scenarios. It ensures that the model has the necessary resources to handle concurrent requests without performance degradation, making it the correct choice for scaling LLM applications beyond simple development or low-traffic testing environments.
- ✗
Disable auto-scaling and fix the number of replicas to one.
Why it's wrong here
Fixing the number of replicas to one eliminates the ability to scale, which is the opposite of what is required for high-concurrency throughput. If traffic spikes, the single replica will become a bottleneck, leading to timeouts and a degraded user experience, failing the scalability requirement.
- ✗
Configure the endpoint to use an external API key for every request.
Why it's wrong here
Using external API keys for every request does not address scaling or throughput concerns. Instead, it adds authentication overhead to every call, which could increase latency. This configuration is related to security, not the architectural optimization needed for handling high volumes of concurrent LLM inference requests.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.