Databricks-ML-Assoc Databricks Machine Learning Practice Question
A data scientist is monitoring model drift in Databricks. Which TWO approaches are recommended to detect performance degradation in a production model?
⚠ Common exam trap
Candidates often focus only on model performance metrics while ignoring the input data. They fail to realize that data drift is a leading indicator of future model performance degradation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Regularly compare current prediction metrics against the baseline training metrics.
Monitoring model drift is essential to ensure the continued accuracy of production models. Comparing current model performance metrics against a baseline established during training allows for proactive retraining. Similarly, tracking the distribution of input features (data drift) helps identify changes in data patterns that may undermine the model's reliability. Combining these methods ensures a comprehensive view of the model's health and readiness for redeployment or adjustment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Regularly compare current prediction metrics against the baseline training metrics.
Why this is correct
Comparing current metrics against baseline metrics is the standard way to detect concept drift. When performance metrics like F1-score or RMSE drop significantly from the training baseline, it serves as a strong indicator that the model no longer accurately represents the current data distribution, necessitating a retraining process.
- ✗
Delete the original training data to force the model to learn new patterns.
Why it's wrong here
Deleting the original training data is destructive and prevents any form of benchmarking or comparative analysis. It makes it impossible to verify the baseline performance or reproduce the original model, which is a violation of basic machine learning management principles. Drift should be addressed by retraining, not data destruction.
- ✓
Monitor feature distributions for statistically significant changes compared to training data.
Why this is correct
Data drift occurs when input data distributions shift over time. Monitoring features using statistical tests (like Kolmogorov-Smirnov) allows engineers to identify these shifts early. Detecting data drift is often a leading indicator of performance degradation, allowing for proactive intervention before the model's output quality drops in production environments.
- ✗
Automate model retraining to run every hour, regardless of drift.
Why it's wrong here
Unnecessary retraining is costly and can lead to model instability if the retraining process is not robust. Retraining should be triggered based on observed drift or schedule, not purely on an hourly basis. A drift-based approach ensures that resources are allocated efficiently and model updates only occur when needed.
- ✗
Manually inspect every prediction in the production logs for errors.
Why it's wrong here
Manual inspection is impossible at scale and prone to human error. Automation is required for effective monitoring. Modern Databricks pipelines should use automated monitoring tools to flag potential drift or quality issues, allowing teams to focus on meaningful model improvements rather than tedious manual review of prediction logs.
About these practice questions
One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.