Databricks-ML-Pro Model Deployment Practice Question
An ML engineer is deploying a model to Databricks Model Serving that uses a custom transformer requiring a GPU. The endpoint must handle high throughput with low latency. Which workload type and configuration should be selected?
⚠ Common exam trap
The trap here is assuming that GPU can be enabled through environment variables or conda specifications, when in fact it must be selected as a workload type during endpoint configuration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a 'GPU_SMALL' or 'GPU_MEDIUM' workload type, depending on the model's memory and compute needs.
Databricks Model Serving provides specific GPU-enabled workload types, such as 'GPU_SMALL' and 'GPU_MEDIUM', which include GPU instances. For a model requiring GPU, selecting one of these workload types is necessary. The choice between small and medium depends on the model's memory and compute demands to achieve high throughput and low latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a 'Large' workload type and specify GPU requirements in the model's conda environment.
Why it's wrong here
Specifying GPU requirements in the conda environment does not provision GPUs; the workload type determines the compute resources. The 'Large' workload type is CPU-based and does not include GPUs. To use GPUs, you must select a GPU-enabled workload type, which is not 'Large'.
- ✗
Use a 'Medium' workload type and configure the endpoint to use GPU by setting an environment variable.
Why it's wrong here
The 'Medium' workload type is CPU-only. Setting an environment variable cannot enable GPU resources because the underlying instances lack GPUs. This approach would not satisfy the GPU requirement and would lead to failures or degraded performance for models that depend on GPU acceleration.
- ✗
Use a 'Small' workload type with CPU and enable GPU acceleration via a model parameter.
Why it's wrong here
The 'Small' workload type is CPU-only and does not support GPU acceleration. Enabling GPU via a model parameter is not possible because the underlying infrastructure does not include GPUs. This configuration would fail to meet the GPU requirement and likely result in poor performance for GPU-dependent models.
- ✓
Use a 'GPU_SMALL' or 'GPU_MEDIUM' workload type, depending on the model's memory and compute needs.
Why this is correct
Databricks Model Serving offers GPU-enabled workload types such as 'GPU_SMALL' and 'GPU_MEDIUM'. These provide GPU instances suitable for models that require GPU acceleration. Selecting the appropriate size based on memory and compute requirements ensures high throughput and low latency for GPU-dependent models.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.