Courseiva
← Back to Databricks Certified Machine Learning Professional questions

Scenario-based practice

Hard Difficulty Questions

Practise Databricks Certified Machine Learning Professional practice questions — original exam-style scenarios covering every exam domain, with detailed explanations, wrong-answer analysis, and common exam traps.

20
scenario questions
Databricks-ML-Pro
exam code
Databricks
vendor

Scenario guide

How to approach hard difficulty questions

These are the questions most candidates get wrong. They require connecting multiple concepts, reading tricky output, or knowing edge-case behaviour that isn't on most study cards. Practising them trains you to operate under uncertainty — a necessary skill on the real exam.

Quick answer

Hard Difficulty Questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Related practice questions

Related Databricks-ML-Pro topic practice pages

Scenario questions usually connect to one or more exam topics. Use these links to review the underlying concepts behind the scenario.

Practice set

Practice scenarios

Question 1hardmultiple choice
Full question →

A team notices that their model performance is significantly lower in production than in training. They suspect 'data drift' in the feature inputs. Which Databricks capability should be used to monitor this?

Question 2hardmultiple choice
Full question →

When using MLflow to manage the machine learning lifecycle, what is the primary purpose of the 'conda.yaml' or 'requirements.txt' file automatically generated during log_model?

Question 3hardmultiple choice
Full question →

A team is building an automated retraining pipeline. They need to ensure that only models exceeding a certain performance threshold are registered. What is the most effective way to implement this logic?

Question 4hardmultiple choice
Full question →

A machine learning engineer is developing a scikit-learn model on Databricks and wants to ensure that the model's input schema is enforced during scoring. They log the model using `mlflow.sklearn.log_model` with the `signature` parameter. Which statement accurately describes the behavior when the model is later served via MLflow Model Serving?

Question 5hardmultiple choice
Full question →

Refer to the exhibit. A machine learning team has updated their model serving endpoint configuration as shown in the JSON. Which deployment strategy is being implemented, and what is the primary risk associated with this specific configuration?

Exhibit

{
  "served_entities": [
    {
      "name": "churn-model-v1",
      "entity_name": "prod.ml_models.churn_prediction",
      "entity_version": "1",
      "workload_size": "Small",
      "scale_to_zero_enabled": true,
      "traffic_config": {
        "percent": 90
      }
    },
    {
      "name": "churn-model-v2",
      "entity_name": "prod.ml_models.churn_prediction",
      "entity_version": "2",
      "workload_size": "Small",
      "scale_to_zero_enabled": true,
      "traffic_config": {
        "percent": 10
      }
    }
  ]
}
Question 6hardmultiple choice
Full question →

A machine learning engineer is using MLflow to track experiments on Databricks. They notice that when they run `mlflow.log_artifact` with a local file path inside a notebook, the artifact is stored in the run's artifact location, but when they run the same code in a job cluster, the artifact is missing. The job cluster uses the same MLflow tracking server and experiment. What is the most likely reason for the missing artifact?

Question 7hardmultiple choice
Full question →

A fraud detection model is served on a Databricks Model Serving endpoint. You notice that predictions for the same input vector differ between two consecutive requests within seconds, and there is no feature store or external cache involved. The model was logged with a fixed random seed and deterministic inference code. Which action should you take first to diagnose the inconsistency?

Question 8hardmultiple choice
Full question →

When auditing an ML pipeline in Databricks for compliance and governance, which THREE of the following should be verified?

Question 9hardmultiple choice
Full question →

Your organization is implementing an MLOps strategy that requires strict model governance. You need to ensure that no model is deployed to production unless it has been tagged with 'validated=true' in the MLflow Model Registry. How can you enforce this policy within your CI/CD workflow?

Question 10hardmulti select
Full question →

A team is using Databricks Feature Store to manage features for a real-time model served via Databricks Model Serving. They need to ensure that the online feature values used at inference time are consistent with the training data. Which TWO practices should they implement? (Choose two.)

Question 11hardmultiple choice
Full question →

A team deploys a model to a Databricks Model Serving endpoint and enables Inference Tables. They later notice that requests are being served successfully but no rows are appearing in the inference table. Which explanation is most likely?

Question 12hardmultiple choice
Full question →

Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?

Exhibit

GET /serving-endpoints/customer-churn/config
{
  "name": "customer-churn",
  "config": {
    "served_entities": [
      {
        "entity_name": "prod.models.churn",
        "entity_version": "5",
        "scale_to_zero_enabled": false
      }
    ]
  }
}
Question 13hardmultiple choice
Full question →

Which THREE of the following are primary responsibilities of an MLOps engineer when maintaining production ML models in Databricks?

Question 14hardmultiple choice
Full question →

A financial institution deploys a credit scoring model using Databricks Model Serving. The model must log all incoming requests and outgoing responses to a Delta table for auditing. The ML engineer needs to enable this logging with minimal performance impact. Which solution should they implement?

Question 15hardmulti select
Full question →

You are auditing a Databricks environment to ensure compliance. Which TWO actions ensure the highest level of model lineage and reproducibility for models registered in MLflow?

Question 16hardmulti select
Full question →

A data science team is transitioning from batch prediction to real-time serving using MLflow. They need to ensure that the deployment process is reproducible and allows for easy rollback. Which TWO actions should the team take to meet these requirements?

Question 17hardmultiple choice
Review the full routing breakdown →

An ML engineer is deploying a model to Databricks Model Serving and wants to implement A/B testing between two model versions. The engineer needs to route a percentage of traffic to each version and collect performance metrics. Which feature of Databricks Model Serving should the engineer use?

Question 18hardmultiple choice
Full question →

A machine learning engineer is using MLflow Tracking to log metrics and artifacts for a deep learning model. They notice that the training run logs a large number of metrics (e.g., loss per batch) and want to reduce the storage footprint and improve query performance. They also need to retain the ability to compare runs and reproduce results. Which of the following actions is most appropriate?

Question 19hardmulti select
Full question →

Which THREE actions are required when migrating a custom Scikit-learn model to the Databricks Model Registry to ensure it can be served via MLflow Model Serving?

Question 20hardmultiple choice
Full question →

Refer to the exhibit. The model serving endpoint is failing during high-latency requests. What is the most appropriate configuration change to resolve this?

Exhibit

LOG_LEVEL: INFO
MLFLOW_TRACKING_URI: databricks
DEPLOY_ENVIRONMENT: PROD
ENABLE_AUTO_REGISTRATION: TRUE
SCORING_TIMEOUT: 60s
ERROR: RequestTimeoutException: Model failed to respond within 60s

These Databricks-ML-Pro practice questions are part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style Databricks-ML-Pro questions with detailed explanations, topic-based practice, mock exams, readiness tracking, and study analytics.