Courseiva
Application Development →hardMultiple Choice

Databricks-GenAI-Assoc Application Development Practice Question

An AI engineer is deploying a RAG application using Databricks Model Serving. They need to ensure the endpoint can handle high traffic with low latency and automatically scale based on demand. Which configuration should they use?

⚠ Common exam trap

The trap here is focusing only on scale-to-zero for cost savings without considering the need for sufficient concurrency and autoscaling to handle high traffic.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Configure autoscaling with a minimum of 1 and maximum of 10 replicas, and set appropriate concurrency per replica.

Autoscaling with a minimum and maximum replica count allows the endpoint to scale out during high traffic and scale in during low traffic, optimizing both latency and cost. Setting appropriate concurrency per replica ensures efficient utilization. This configuration provides the elasticity required for high traffic with low latency, while avoiding the pitfalls of fixed provisioning or overly restrictive concurrency limits.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Configure autoscaling with a minimum of 1 and maximum of 10 replicas, and set appropriate concurrency per replica.

    Why this is correct

    Autoscaling dynamically adjusts the number of replicas based on load, ensuring low latency during traffic spikes and cost savings during idle periods. Setting a minimum of 1 avoids cold starts, while a maximum of 10 caps resource usage. Concurrency per replica should be tuned to the model's throughput. This meets the requirements for high traffic and automatic scaling.

  • ✗

    Enable scale-to-zero and set a maximum concurrency of 1.

    Why it's wrong here

    Scale-to-zero reduces cost when idle, but setting maximum concurrency to 1 severely limits throughput. Each request would be processed sequentially, causing high latency under load. This configuration is unsuitable for high traffic; it would result in queuing and timeouts. Concurrency should be set based on the model's capacity and expected load.

  • ✗

    Disable autoscaling and manually provision 20 replicas to handle peak load.

    Why it's wrong here

    Manually provisioning for peak load ensures capacity but wastes resources during low traffic, increasing cost. It also lacks automatic adjustment if load exceeds 20 replicas. Autoscaling is designed to balance performance and cost by scaling in and out as needed, which is preferable for dynamic workloads.

  • ✗

    Use a single large replica with maximum concurrency set to 100.

    Why it's wrong here

    A single replica, even with high concurrency, has finite resources. Under high traffic, it may become a bottleneck, leading to increased latency and potential failures. Without autoscaling, there is no ability to add capacity during spikes. This approach lacks the elasticity needed for varying demand.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.