Databricks-ML-Pro ML Ops Practice Question
When auditing an ML pipeline in Databricks for compliance and governance, which THREE of the following should be verified?
⚠ Common exam trap
Candidates often include irrelevant metrics like model accuracy or latency as audit requirements, failing to realize that compliance audits focus on traceability, lineage, and access control.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The Git commit hash associated with the specific training run.
Auditing requires proof of data lineage, code versioning, and access control. Verifying these components ensures that the organization can prove how a model was built, what data was used, who approved it, and that the code was properly reviewed. This level of traceability is non-negotiable for regulated industries and essential for enterprise-grade MLOps maturity.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The Git commit hash associated with the specific training run.
Why this is correct
Traceability starts with code. By linking every model to a specific Git commit, auditors can verify the exact code state that produced the model. This is fundamental for reproducibility and ensuring that no unauthorized or unreviewed changes were introduced into the model generation process.
- ✗
The personal credentials of the data scientist who manually ran the training.
Why it's wrong here
Individual user credentials should never be used for production pipelines. Pipelines should be executed by service principals to ensure consistency, security, and traceability. Relying on personal credentials creates a single point of failure and makes the audit trail dependent on individual employees rather than organizational roles.
- ✓
The lineage of the training data, including source table versions.
Why this is correct
Data lineage is critical for understanding what informed the model. Without knowing the exact version of the training data, reproducing a model's performance is impossible. Auditors need to confirm that training data was sourced from valid, versioned tables to ensure the model was built correctly.
- ✓
The list of authorized users who can transition models to 'Production'.
Why this is correct
Controlling the model promotion process is key to governance. Verifying that only authorized users or automated CI/CD service principals can move a model to 'Production' ensures that no unvetted models reach end-users, reducing the risk of security vulnerabilities or deployment of poor-quality machine learning models.
- ✗
The raw text of every email sent between the data scientists.
Why it's wrong here
Communicating via email for technical decisions is not a verifiable part of a pipeline audit. Compliance relies on logs, metadata, and automated records stored within the Databricks platform. Email history is irrelevant and impossible to track in a systematic, automated audit of the ML model lifecycle.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.