Courseiva
Model Deployment →hardMultiple Select

Databricks-ML-Assoc Model Deployment Practice Question

Which THREE factors should be considered when choosing the 'workload size' (e.g., Small, Medium, Large) for a Databricks Model Serving endpoint?

⚠ Common exam trap

Candidates often select cost or model accuracy instead of operational metrics like memory footprint, traffic volume, and latency targets when sizing endpoints.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The memory requirements of the model artifact

Selecting the correct workload size is a trade-off between latency requirements, compute costs, and the model's resource footprint. Larger models with high parameter counts often require 'Large' configurations to prevent memory exhaustion or high tail latency. Conversely, smaller models may perform adequately on 'Small' instances, which are cost-effective. Monitoring the endpoint's resource utilization metrics is key to right-sizing the deployment and ensuring optimal performance within the budget constraints.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The memory requirements of the model artifact

    Why this is correct

    The model artifact size and its runtime memory consumption directly dictate the minimum compute resources needed. If the model requires more RAM than the chosen workload size provides, the endpoint will fail to load or experience frequent crashes. Assessing memory usage during the testing phase is critical for size selection.

  • ✗

    The total number of users in the workspace

    Why it's wrong here

    The number of users in the workspace is generally irrelevant to the serving endpoint configuration. The workload size is determined by the inference request volume and the computational demands of the model itself. Scaling should be based on incoming traffic patterns rather than the number of developers in the workspace.

  • ✓

    Expected request latency targets

    Why this is correct

    Different workload sizes provide varying CPU and memory resources, which directly impact the inference latency. If a business requirement demands low latency for a model that is computationally heavy, selecting a larger workload size can help meet these targets by providing more compute power for the inference operation.

  • ✓

    The volume of expected incoming traffic

    Why this is correct

    High request throughput necessitates higher compute allocation to ensure the queue does not grow and cause latency spikes. The workload size defines the capacity per container, and autoscaling will then manage the number of replicas. Choosing the right size ensures that each replica handles its share of traffic efficiently.

  • ✗

    The color scheme of the MLflow UI

    Why it's wrong here

    User interface settings in Databricks are purely cosmetic and have zero impact on the backend infrastructure or the performance of a deployed model. These UI preferences should never influence architectural decisions regarding compute provisioning, as they are completely decoupled from the model serving engine and its operational requirements.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.