Databricks-GenAI-Assoc Application Development Practice Question
Which TWO actions are necessary to ensure that a model serving endpoint in Databricks remains available and performant during peak traffic hours?
⚠ Common exam trap
Students often select only one correct action or mistakenly choose manual scaling scripts, forgetting that both auto-scaling and GPU acceleration are required for high performance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable auto-scaling for the serving endpoint.
Maintaining performance requires both vertical and horizontal scaling strategies. First, configuring auto-scaling on the endpoint allows the infrastructure to adjust to fluctuating demand automatically. Second, ensuring the endpoint uses appropriate instance types—such as GPU-accelerated instances for LLMs—ensures the compute capacity is sufficient for inference tasks. These two actions are foundational for achieving high availability and low latency, preventing request timeouts and ensuring a smooth experience during heavy usage periods.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Disable the serving endpoint's logging to save on compute cycles.
Why it's wrong here
Disabling logging does not meaningfully impact the inference performance of the serving endpoint, as logging is an asynchronous process. Furthermore, turning off observability prevents developers from diagnosing bottlenecks during peak load, which is counterproductive when trying to maintain and optimize the performance of the model serving infrastructure.
- ✓
Enable auto-scaling for the serving endpoint.
Why this is correct
Auto-scaling allows the serving endpoint to dynamically adjust its instance count based on current request volume. This ensures that the system handles spikes in traffic effectively without manual intervention, maintaining consistent performance and avoiding service degradation when user demand increases suddenly during peak hours or application usage cycles.
- ✗
Configure the endpoint to run on a single, fixed-size node to minimize latency.
Why it's wrong here
A single, fixed-size node is prone to becoming a bottleneck during peak traffic, leading to increased latency and potential service outages. Scaling horizontally across multiple nodes is the best practice for high availability, allowing the workload to be distributed effectively and ensuring resilience against high traffic demands.
- ✓
Select appropriate GPU-accelerated compute types for the workload.
Why this is correct
LLM inference is computationally intensive and benefits significantly from hardware acceleration. Selecting GPU-based instances specifically designed for AI workloads provides the necessary performance to meet low-latency requirements, ensuring that the model can process requests rapidly even under high concurrent load, which is critical for a positive user experience.
- ✗
Manually restart the endpoint every hour to clear the cache.
Why it's wrong here
Restarting an endpoint causes unnecessary downtime and flushes the model cache, which can actually increase latency for subsequent requests as the model reloads. This is an anti-pattern that disrupts availability and does not address the underlying need for scalable infrastructure, which should be managed by the platform automatically.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.