Courseiva

Databricks Certified Machine Learning Professional (Databricks-ML-Pro) — Questions 1–75

300 questions total · 4pages · All types, answers revealed

Page 1 of 4

Page 2
1
MCQeasy

A data scientist is using MLflow on Databricks to track a series of experiments. They want to compare the performance of different runs and identify the best model based on a custom metric called "f1_score". Which MLflow feature should they use to efficiently compare and rank these runs?

A.MLflow Experiments page in the Databricks workspace, using the metric column and sorting.
B.MLflow Model Registry, by registering each run's model and comparing versions.
C.Databricks SQL, by creating a dashboard that queries the MLflow tracking database directly.
D.MLflow Tracking API, by writing a custom script to query the REST API and compute rankings.
AnswerA

The MLflow Experiments UI allows users to view all runs in an experiment, display metrics as columns, and sort by any metric. This makes it easy to compare runs and identify the one with the highest f1_score. It is the primary tool for run comparison and requires no additional code.

Why this answer

The MLflow Experiments page in Databricks provides a built-in, user-friendly interface to view all runs, display metrics, and sort by any metric. This allows quick identification of the best run based on f1_score. Other options either require unnecessary coding, are not designed for run comparison, or use unsupported methods.

Exam trap

The trap here is overcomplicating the solution by considering custom scripts or direct database queries when the UI already provides the needed functionality.

2
MCQhard

A data scientist is using MLflow to log a model that includes a custom preprocessing step. They want to ensure that the preprocessing is applied consistently during both training and inference. Which approach should they take?

A.Use `mlflow.sklearn.log_model` and include the preprocessing steps in a scikit-learn `Pipeline` object.
B.Log the preprocessing code as a separate artifact and manually apply it before calling the model during inference.
C.Log the model with `mlflow.pyfunc.log_model` and provide the preprocessing function as a separate file in the `artifacts` parameter.
D.Create a custom `PythonModel` that includes both the preprocessing and the model, and log it with `mlflow.pyfunc.log_model`.
AnswerD

A custom `PythonModel` can encapsulate both preprocessing and the model's prediction logic. When logged with `mlflow.pyfunc.log_model`, the entire pipeline is packaged as a single model. During inference, MLflow loads this model and applies the preprocessing automatically, ensuring consistency between training and inference.

Why this answer

Creating a custom `PythonModel` that integrates preprocessing and prediction ensures that the entire pipeline is encapsulated in a single MLflow model. When logged, this model applies preprocessing automatically during inference, guaranteeing consistency. This is the most robust approach when preprocessing involves custom logic not easily represented in a standard pipeline.

Exam trap

The trap here is assuming that logging preprocessing as an artifact or using a scikit-learn Pipeline always works; custom logic may require a PythonModel for full encapsulation.

3
MCQhard

Your organization wants to monitor production models for drift. Which Databricks service should be used to detect changes in the input data distribution compared to the training data?

A.MLflow Model Registry
B.Databricks Lakehouse Monitoring
C.Delta Live Tables (DLT)
D.Databricks SQL Alerting
AnswerB

Lakehouse Monitoring is specifically built to compute drift metrics by comparing production data with training baselines. It provides automated alerts and dashboards, enabling teams to proactively identify when input feature distributions shift, which is a primary cause of model performance degradation in production environments.

Why this answer

Databricks Lakehouse Monitoring is the native solution for tracking data quality and drift. By analyzing incoming data against a baseline established during training, it identifies statistical deviations. This is critical for MLOps because detecting drift early allows data scientists to trigger retraining or investigate data pipeline issues, ensuring model performance remains consistent over time despite changing real-world data patterns.

Exam trap

Test-takers frequently look for manual logging solutions, missing the native Databricks Lakehouse Monitoring service designed explicitly for automated drift detection.

4
MCQeasy

A data scientist is using MLflow Tracking to log experiments for a model that predicts customer lifetime value. They want to compare runs across multiple experiments and identify the best performing run based on a custom metric called 'rmse'. Which MLflow feature should they use?

A.MLflow Models
B.MLflow Model Registry
C.MLflow Projects
D.MLflow Tracking UI and API
AnswerD

The MLflow Tracking UI allows users to visualize and compare runs across experiments, including metrics like 'rmse'. The API (e.g., MlflowClient.search_runs) enables programmatic querying and filtering based on metrics. This is the correct tool for comparing runs and identifying the best one.

Why this answer

MLflow Tracking provides both a UI and an API to log, query, and compare runs. The UI allows visual comparison of metrics like 'rmse' across experiments, and the API supports programmatic search and filtering. This makes it the appropriate feature for identifying the best performing run based on a custom metric.

Exam trap

The trap here is confusing MLflow Tracking with MLflow Model Registry, which manages model lifecycle but not run comparison.

5
MCQmedium

Refer to the exhibit. An engineer is configuring a serving endpoint. Based on the configuration provided, what is the impact of the 'auto_scale' flag?

A.It scales the number of model versions deployed in the registry.
B.It allows the serving endpoint to adjust worker count based on request load.
C.It automatically re-trains the model when request volume drops.
D.It forces the endpoint to only use 10 workers permanently.
AnswerB

Auto-scaling is specifically designed to manage infrastructure resources dynamically. By adjusting the number of workers based on demand, the endpoint remains responsive while avoiding the unnecessary cost of provisioning for maximum load at all times, making it a best practice for production deployments.

Why this answer

The 'auto_scale' flag allows the Databricks Model Serving endpoint to dynamically adjust the number of active workers based on real-time traffic load. This ensures that the system maintains performance during peak request periods while minimizing costs during low-traffic periods. This balance of performance and efficiency is a critical aspect of managing production ML infrastructure in an enterprise Databricks environment.

Exam trap

Candidates often assume auto_scale refers to cluster resizing or data processing parallelism, failing to realize it specifically governs the number of compute instances serving the model endpoint based on traffic.

6
MCQhard

A fraud-detection model is registered in Unity Catalog and deployed to a Model Serving endpoint. Compliance requires that every prediction be traceable to the exact model version and the request that produced it. Which combination of Databricks features should you configure to meet this requirement?

A.Enable inference table logging on the endpoint and register the model version in Unity Catalog with a descriptive comment.
B.Attach a model signature to the registered model and enable request logging in the serving client application.
C.Configure the endpoint with a scale-to-zero policy and log the endpoint configuration to a Delta table daily.
D.Enable MLflow autologging in the training notebook and set the model version stage to Production.
AnswerA

Inference tables record each request, response, timestamp, and the served model version, while Unity Catalog provides the lineage and version identity of the registered model. Together they let an auditor trace any prediction back to the exact model version and the originating request, which is precisely the compliance requirement.

Why this answer

Traceability of individual predictions requires server-side capture of the request, response, and serving model version, which is what inference tables provide, combined with Unity Catalog's governed model version identity. Training-time logging, endpoint configuration snapshots, and client-side logs each capture only part of the picture and cannot reconstruct the full per-prediction audit trail.

Exam trap

The trap here is assuming that MLflow autologging or a model signature provides runtime prediction traceability, when those features operate at training time or schema validation only.

7
MCQmedium

When using the Databricks Model Registry, what does a 'Model Version' represent in the context of the lifecycle?

A.An automated report generated after every training run.
B.A unique, immutable snapshot of a specific model artifact and its metadata.
C.A live pointer that always points to the most recent training run.
D.A shared directory containing the raw training data and configuration files.
AnswerB

Each version in the Model Registry is immutable, ensuring that once a model is registered, it cannot be tampered with. This immutability is critical for compliance and reproducibility, as it guarantees that the exact model code and weights used in training are exactly what gets served in production.

Why this answer

A Model Version is a point-in-time snapshot of a model, including the code, model weights, dependencies, and environment configuration. By tracking these versions, data scientists can compare performance across different iterations, rollback to previous states if a production deployment fails, and maintain an audit trail of how a model evolved over time. This versioning is foundational for reproducible machine learning.

Exam trap

Candidates often confuse a 'Model Version' with the 'Registered Model' (the container) or a 'Run' (the training instance), failing to distinguish the immutable artifact snapshot.

8
MCQhard

Refer to the exhibit. The JSON configuration represents an existing Databricks Model Serving endpoint. You need to update this endpoint to support a traffic split between version 5 and version 6 for A/B testing. Which update strategy is correct?

A.Update the config to replace 'model_version': '5' with 'model_version': '6' in the current served_models list.
B.Create two separate endpoints, one for version 5 and one for version 6, then split traffic using a load balancer.
C.Modify the served_models array to include both versions with specific traffic weights assigned to each.
D.Delete the current endpoint and recreate it with a new configuration that includes only version 6.
AnswerC

Adding both versions to the `served_models` list enables traffic routing control. By specifying the `traffic` field for each, you can define the percentage of requests allocated to each version. This configuration is the standard method for A/B testing in Databricks, providing safe, granular control over model deployment transitions.

Why this answer

To perform A/B testing, you must define multiple served models within the `served_models` array in the endpoint configuration. Each entry requires a `traffic` weight that sums to 100%. This allows Databricks to route incoming requests according to the specified percentages.

This method is crucial for safely rolling out new models, as it allows you to observe performance metrics on a subset of real-world traffic before a full production cutover.

Exam trap

Candidates try to create multiple endpoints for A/B testing rather than modifying the `served_models` array within a single endpoint configuration to split traffic weights.

9
Multi-Selectmedium

A team is preparing to promote a new model version to production in the MLflow Model Registry. They must ensure the model can be served with a consistent environment across staging and production and that dependency drift is detected before promotion. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Store only the model's pickle file and reconstruct dependencies manually on each environment, since pickle serialization is environment-independent.
B.Log the model with its conda environment and requirements files so the exact dependency versions are captured as part of the model version.
C.Rely on the serving endpoint to install the latest available versions of each library at deployment time so the environment stays current.
D.Verify the logged environment files against the target serving environment before promotion, resolving any version mismatches in the model's dependency specification.
E.Pin the serving cluster's runtime to a different Databricks Runtime version than training used, to validate cross-version compatibility during staging.
AnswersB, D

Logging the conda environment and requirements files records the exact library versions used at training time as part of the model artifact. When the same model version is served in staging and production, the environment can be reconstructed from these files, which is the foundation for consistent, reproducible serving and for detecting any drift from the recorded versions.

Why this answer

Capturing the exact dependency environment with the model version and then validating that environment against the target serving environment before promotion together ensure consistent serving and catch drift early. Floating to latest versions, mismatching runtimes, or relying on bare pickles all break reproducibility or hide drift.

Exam trap

The trap here is thinking that a serialized model file alone guarantees reproducible serving, when the surrounding dependency environment must also be captured and verified.

10
MCQmedium

Your team is experiencing 'training-serving skew' where model performance in production is significantly lower than during training. Which approach should you prioritize to mitigate this issue?

A.Increase the amount of historical training data used in the model.
B.Use the Databricks Feature Store to unify feature engineering pipelines.
C.Switch the model architecture to a deeper neural network.
D.Retrain the model every hour to keep it fresh with recent data.
AnswerB

The Feature Store enforces consistent data processing by design. By using the same feature definitions for both training and serving, you ensure that the input features are computed identically in both environments, which is the industry-standard method for resolving training-serving skew in production machine learning.

Why this answer

Training-serving skew often occurs because the data transformation logic used during model training differs from the logic applied during real-time inference. By utilizing the Databricks Feature Store, you can centralize your feature engineering code. This ensures that the exact same transformation code and look-up logic are utilized for both offline training and online inference, effectively eliminating the discrepancy between the two environments.

Exam trap

Candidates often look for retraining frequency adjustments or hyperparameter tuning, failing to recognize that training-serving skew is primarily caused by divergent feature pipelines.

11
Multi-Selectmedium

A machine learning team is using Databricks to develop a model and wants to ensure that the model's input schema is validated at inference time to prevent errors from malformed data. Which TWO approaches allow them to enforce schema validation when serving the model with MLflow Model Serving? (Choose two.)

Select 2 answers
A.Enable MLflow Model Registry webhooks to validate the model signature before each inference request.
B.Log the model with an MLflow signature that specifies the expected input schema, and rely on MLflow Model Serving to reject requests that do not conform to the signature.
C.Use Databricks Feature Store to define the feature schema and rely on it to validate incoming requests at serving time.
D.Configure Model Serving to use a Delta table as the input source and enable schema evolution, so that any schema changes are automatically handled.
E.Implement a custom Python function that checks the input schema and raises an exception if it is invalid, then log the model using `mlflow.pyfunc.log_model` with that function as the model's `predict` method.
AnswersB, E

When an MLflow model is logged with a signature, Model Serving uses it to validate incoming requests. If the payload does not match the schema (e.g., missing columns, wrong data types), the request is rejected with an error. This provides automatic schema enforcement without additional code. It is the recommended way to ensure input validation for served models.

Why this answer

Logging a model with an MLflow signature enables Model Serving to automatically validate incoming requests against the expected schema, rejecting mismatches. Alternatively, a custom `pyfunc` model can implement explicit validation logic in its `predict` method. Both approaches enforce schema validation at inference time.

The other options do not provide runtime validation for served models.

Exam trap

The trap here is assuming that tools like Feature Store or Model Registry webhooks can validate inference requests, when they actually operate at training or registry-event time, not at serving time.

12
MCQmedium

Which component of MLflow is responsible for keeping track of the different versions of a model as it moves from development to testing and production?

A.MLflow Tracking
B.MLflow Projects
C.MLflow Model Registry
D.MLflow Recipes
AnswerC

The Model Registry provides a centralized store for managing the full lifecycle of MLflow models. It supports versioning, stage transitions, and annotations. It is the primary tool in Databricks for managing production deployments, ensuring that the correct model versions are used in the correct environments at all times.

Why this answer

The MLflow Model Registry is the central repository for model versioning. It allows teams to manage the lifecycle of a model by transitioning versions through different stages. This is a critical MLOps function because it ensures that production environments are always pointing to a known, stable version of the model, while allowing developers to continue iterating on new versions without disrupting the live service.

Exam trap

Candidates often confuse the MLflow Tracking Server with the Model Registry, failing to distinguish between experiment logging and the centralized management of model production versions.

13
Multi-Selecthard

Your team uses Databricks for machine learning and needs to ensure that model training is fully automated and reproducible. Which THREE of the following are necessary components for a production-grade automated ML pipeline?

Select 3 answers
A.Git integration to version control the training notebooks and scripts.
B.Databricks Workflows to schedule and orchestrate pipeline steps.
C.MLflow Model Registry to manage model versions and deployment lifecycle.
D.Manual approval steps for every single model training run.
E.Using the same cluster for both development and production tasks to minimize complexity.
AnswersA, B, C

Version control is fundamental to MLOps. It allows teams to track changes, collaborate, and revert to known good states. Without Git, it is impossible to audit the evolution of training logic or ensure that the code running in production is exactly the same as what was tested during the CI process.

Why this answer

An automated pipeline requires a workflow orchestrator (Workflows), a centralized registry (MLflow), and version-controlled logic (Git). These three components work together to provide a seamless transition from code development to production deployment. This triad ensures that every step is reproducible, monitored, and audit-ready, which is the foundational requirement for any mature MLOps practice operating at enterprise scale.

Exam trap

Candidates often include 'manual monitoring' or 'local testing' as key components, failing to identify the three pillars of automation: Git, Workflows, and the Model Registry.

14
MCQhard

You are monitoring a model served on Databricks Model Serving. You need to detect data drift in the incoming requests without delaying predictions. Which approach should you use?

A.Add a pre-processing step in the model's predict function that computes drift metrics on each request.
B.Configure the endpoint to call an external monitoring service synchronously before returning predictions.
C.Use Databricks SQL dashboards to query the model's training data and compare it to live predictions in real time.
D.Enable inference logging on the endpoint and analyze the logged requests asynchronously using a Databricks job that computes drift metrics.
AnswerD

Inference logging captures the request payloads and predictions to a Delta table. You can then run an asynchronous job to compute drift metrics, such as population stability index or KL divergence, comparing recent data to a baseline. This does not add latency to predictions because logging is decoupled from the serving path, making it the recommended approach.

Why this answer

Inference logging on Model Serving captures request and response data to a Delta table without adding latency to the prediction path. An asynchronous job can then analyze this logged data to compute drift metrics, enabling detection without impacting performance. Synchronous or in-model computations are impractical due to latency and single-request limitations.

Exam trap

The trap here is assuming that drift detection must happen inline with predictions, but it should be decoupled to avoid latency and allow windowed analysis.

15
Multi-Selectmedium

A team is deploying a model to Databricks Model Serving and wants to implement a canary release strategy to gradually shift traffic from the current model version to a new version. Which TWO configurations are required to achieve this? (Choose two.)

Select 2 answers
A.Create a serving endpoint with multiple served entities, each referencing a different model version.
B.Deploy two separate endpoints and use an external load balancer to distribute traffic between them.
C.Enable automatic canary deployment by setting a flag in the model's MLflow metadata.
D.Configure the endpoint to use a single served entity and rely on the model's internal logic to route requests based on a random seed.
E.Use the Databricks REST API to update the traffic configuration for the endpoint, specifying the percentage of traffic for each served entity.
AnswersA, E

Databricks Model Serving supports multiple served entities within a single endpoint, each pointing to a different model version. This allows you to route traffic to different versions, which is essential for canary releases. You can then adjust the traffic split between the entities to gradually shift traffic.

Why this answer

Canary releases in Databricks Model Serving are achieved by configuring a single endpoint with multiple served entities, each pointing to a different model version. Traffic is then split between these entities using the endpoint's traffic configuration, which can be updated via the REST API. This allows gradual shifts and monitoring.

Exam trap

The trap here is thinking that canary deployments require multiple endpoints or automatic flags, when actually they are configured within a single endpoint using multiple served entities and explicit traffic weights.

16
MCQhard

A data scientist is using MLflow to track a deep learning experiment on Databricks. They want to log custom metrics that are computed during training but not automatically captured by `mlflow.autolog()`. What is the correct way to log these custom metrics?

A.Use `mlflow.set_tag` to record the metric values.
B.Use `mlflow.log_param` to log the metric values.
C.Use `mlflow.log_metric` within the training loop.
D.Use `mlflow.log_artifact` to save a file containing the metrics.
AnswerC

`mlflow.log_metric` allows logging custom metrics at any point during training. By calling it within the training loop, the data scientist can record metrics such as custom loss functions or evaluation scores that autolog does not capture. This provides flexibility and ensures all relevant metrics are tracked.

Why this answer

The correct method is to use `mlflow.log_metric` within the training loop. This function is designed for logging metrics and supports step-wise logging, which is ideal for tracking custom metrics over epochs. It integrates with MLflow's UI for visualization and comparison.

Exam trap

The trap here is confusing the different logging functions; metrics must be logged with `log_metric` to be properly tracked and visualized.

17
MCQeasy

When developing a machine learning pipeline on Databricks, which feature provides the most effective way to track the lineage of a model from the raw data used for training to the final deployment?

A.Manually maintaining a spreadsheet documenting training data versions.
B.Using MLflow Tracking integrated with Unity Catalog lineage.
C.Storing training data in a local folder on the driver node.
D.Creating a new database for every model training run.
AnswerB

MLflow Tracking captures the training process parameters and metrics, while Unity Catalog tracks the data lineage of the inputs. This combined approach provides a comprehensive view of the model's history, from the raw data source to the final model artifact, meeting strict audit and governance requirements in production.

Why this answer

MLflow Tracking and the Unity Catalog integration are the cornerstones of model lineage in Databricks. Tracking allows developers to log parameters, code versions, and data snapshots, while Unity Catalog provides governance and data lineage. Together, these tools ensure full traceability, which is a mandatory requirement for compliance and auditing in enterprise-grade machine learning systems where understanding how a model reached its current state is critical.

Exam trap

Candidates often rely solely on MLflow tracking for data lineage, forgetting that Unity Catalog is specifically required for end-to-end data governance and dataset lineage tracking.

18
MCQmedium

Refer to the exhibit. You are managing the 'revenue_forecast' model in the registry. A colleague wants to deploy this version to production. Which Databricks command or process is required to move version 4 to the 'Production' stage while ensuring existing production models remain unaffected?

A.Delete the existing Production version and then promote version 4 to Production.
B.Use the MLflow transition_model_version_stage API with archive_existing_versions=True.
C.Directly update the 'stage' tag in the model JSON definition to 'Production'.
D.Reregister the model as a new model name to avoid conflicts with version 3.
AnswerB

This method is the programmatic way to promote a new version while safely archiving the previous one. Setting the parameter to True ensures that the registry remains clean and that there is no ambiguity regarding which model version is currently serving production traffic for the 'revenue_forecast' application.

Why this answer

In Databricks, using the 'transition_model_version_stage' method with 'archive_existing_versions=True' is the standard practice for seamless deployments. This automatically moves the previous production version to 'Archived', maintaining a clear audit trail. Proper lifecycle management ensures that only one version is active in production at a time, preventing conflicts and ensuring that the most current, verified model is the one serving live traffic.

19
MCQmedium

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package not available in the default environment. The model was logged with MLflow and includes the package in its conda environment. What must the engineer ensure for the endpoint to successfully load the model?

A.The package must be installed on the driver node of the Databricks cluster used for serving.
B.The package must be installed via an init script that runs when the endpoint starts.
C.The package must be uploaded to DBFS and referenced in the model's signature.
D.The package must be included in the model's conda environment and the endpoint must be configured to use that environment.
AnswerD

When logging an MLflow model, the conda environment specifies the dependencies. Databricks Model Serving reads this environment file and installs the listed packages into the serving container. The engineer must ensure the custom package is correctly listed in the conda environment and that the endpoint is created without overriding the environment. This allows the serving environment to replicate the training environment, making the custom package available.

Why this answer

Databricks Model Serving builds the serving environment from the MLflow model's conda environment. To use a custom package, it must be listed in that environment. The endpoint will then install it automatically.

Cluster-based methods, DBFS uploads, or init scripts do not affect the serverless serving container, so they are not correct.

Exam trap

The trap here is assuming that serving endpoints run on Databricks clusters where you can install packages or run init scripts, when they actually use isolated serverless containers built from the model's conda environment.

20
MCQhard

A data scientist is using MLflow on Databricks to train a model with a custom training loop. They want to log the model so that it can be loaded later with `mlflow.pyfunc.load_model()` and used for batch inference. The model artifacts include a Python class and a configuration file. Which approach should they use to log the model?

A.Use `mlflow.pyfunc.log_model()` with a custom PythonModel class that defines `predict()` and `load_context()`.
B.Use `mlflow.sklearn.log_model()` with the `signature` parameter.
C.Use `mlflow.tensorflow.log_model()` with `saved_model=True`.
D.Use `mlflow.log_artifact()` to save the Python class and configuration file, then log the model with `mlflow.pyfunc.log_model()` without a PythonModel.
AnswerA

A custom PythonModel allows you to encapsulate preprocessing, model loading, and prediction logic in a single artifact. When loaded with pyfunc.load_model, MLflow instantiates the class and calls predict. This is the documented way to log arbitrary Python models and ensures the configuration file is accessible via the context.

Why this answer

To log a custom Python model that can be loaded as a PyFunc, you must define a PythonModel subclass and pass it to mlflow.pyfunc.log_model. This provides the load_context and predict methods that MLflow uses at inference time. Other flavors are tied to specific libraries and cannot handle arbitrary Python classes.

Exam trap

The trap here is thinking that logging artifacts alongside a model is enough for pyfunc loading, when the model must contain a PythonModel implementation to define inference behavior.

21
MCQeasy

A data scientist has trained a scikit-learn model and logged it with MLflow. They now want to register this model in the MLflow Model Registry and transition it to 'Production' to be served via Databricks Model Serving. Which of the following is a prerequisite for registering the model?

A.The model must be logged using the MLflow Tracking API and have a valid run ID.
B.The model must be logged to an MLflow experiment that is associated with a registered model.
C.The model must be logged with a signature.
D.The model must be approved by a workspace admin.
AnswerA

To register a model, you need a model artifact that is associated with a run in MLflow Tracking. The run ID provides the lineage and allows the registry to reference the model's source. Without a valid run, registration cannot proceed.

Why this answer

Registering a model in MLflow Model Registry requires that the model artifact is linked to an MLflow run. The run ID establishes the model's origin and enables versioning. Other options are either optional or incorrect; signatures, experiment associations, and admin approvals are not required for registration.

Exam trap

The trap here is assuming that a model signature or admin approval is needed to register a model, when actually only a valid MLflow run is required.

22
MCQeasy

A team has deployed a model to a Databricks Model Serving endpoint. They want to monitor the endpoint's performance and detect data drift over time. Which Databricks feature should they use to automatically track inference data and compute drift metrics?

A.Inference tables
B.MLflow tracking server
C.Model serving endpoint logs
D.Databricks SQL dashboards
AnswerA

Inference tables automatically capture the request and response payloads for a serving endpoint and store them in a Delta table. This allows you to monitor model performance, detect data drift, and analyze predictions over time using Databricks SQL or notebooks.

Why this answer

Inference tables are a Databricks feature that automatically logs the input and output of a model serving endpoint to a Delta table. This data can then be used to compute drift metrics, monitor model performance, and trigger alerts. It is the built-in solution for capturing inference data for monitoring purposes.

Exam trap

The trap here is confusing operational logs with inference data capture; endpoint logs do not contain the payloads needed for drift detection.

23
MCQeasy

When developing a machine learning model on Databricks, what is the primary benefit of using Feature Store over standard Delta Lake tables for feature management?

A.It provides built-in GPU acceleration for all training jobs.
B.It automates the elimination of data leakage during training.
C.It ensures feature consistency between training and inference.
D.It automatically converts all data types to tensors.
AnswerC

Feature Store maintains consistent definitions and transformations for features, allowing the same logic to be served to both batch training and real-time inference pipelines. This consistency is essential to prevent training-serving skew, ensuring the model performs as expected when it encounters live data in production environments.

Why this answer

Databricks Feature Store provides a unified interface for feature engineering, storage, and retrieval. Its primary advantage is consistency, ensuring that the same feature transformation logic is applied during both training and inference. This eliminates "training-serving skew," a common issue where discrepancies between development and production pipelines lead to degraded model performance in real-world scenarios.

Exam trap

Candidates often focus on the storage speed of Feature Stores, missing the core MLOps value: eliminating training-serving skew by enforcing identical feature logic for both training and inference.

24
MCQmedium

Which of the following describes the correct usage of the MLflow 'log_param' function in a Databricks environment?

A.It should be called after the model is trained to log the final weights.
B.It logs individual hyperparameter values, such as learning rates or tree depth, to the active run.
C.It automatically logs the entire training dataset for lineage tracking.
D.It is only available for custom models and not for Spark MLlib models.
AnswerB

The 'log_param' function is specifically designed to store configuration settings that influence the model training process. Storing these key-value pairs allows for easy filtering, sorting, and comparison of multiple MLflow runs, which is essential for identifying the best hyperparameters during a grid or random search optimization process.

Why this answer

Logging parameters is a best practice for experiment reproducibility. By calling 'log_param' for every hyperparameter, you ensure that every run is documented with the exact settings used. This allows for rigorous comparison between runs.

In a production environment, this metadata is what enables the team to identify which model version is associated with specific hyperparameter configurations, facilitating model auditing and governance.

Exam trap

Candidates confuse 'log_param' with 'log_metric', incorrectly using parameter logging for runtime outputs or evaluation scores that change iteratively during training.

25
MCQeasy

You are monitoring a model deployed to a Databricks Model Serving endpoint. You need to track the distribution of input features and model predictions over time to detect data drift. Which Databricks feature should you use?

A.Inference tables
B.Delta Live Tables
C.MLflow tracking server logs
D.Databricks SQL dashboards
AnswerA

Inference tables automatically capture the input features and model predictions for each request to a Model Serving endpoint. They store the data in a Delta table, enabling you to analyze feature distributions and detect drift over time. This is the built-in solution for monitoring model serving traffic on Databricks.

Why this answer

Inference tables automatically log the input features and predictions for each request to a Model Serving endpoint, storing them in a Delta table. This allows you to monitor feature distributions and detect data drift over time using SQL or dashboards, making it the correct choice for this scenario.

Exam trap

The trap here is confusing training-time tracking (MLflow) with inference-time monitoring; only inference tables capture production request data for drift detection.

26
MCQhard

A machine learning engineer is using Hyperopt with SparkTrials to tune a scikit-learn model on a Databricks cluster. They set max_evals=100 and parallelism=4. After the run, they notice that some trials report a loss of NaN and that the best model selected by Hyperopt has poor performance. What is the most likely reason for the NaN losses?

A.SparkTrials distributes trials across workers, and the NaN is caused by a serialization error when returning the loss from the worker to the driver.
B.The parallelism parameter is set too high, causing race conditions in the shared search space that corrupt the loss values.
C.The objective function returns a NaN when the model fails to converge, and Hyperopt treats NaN as a valid loss, potentially selecting a failed trial as best.
D.The model's hyperparameters include a regularization parameter that, when set too high, causes the coefficients to become NaN, and Hyperopt does not handle this.
AnswerC

Hyperopt does not automatically filter out NaN losses; it compares them numerically, and NaN comparisons can lead to unpredictable selection. If the objective function returns NaN due to convergence failure or invalid parameters, those trials can be incorrectly considered as having a low loss (since NaN comparisons are false, it may not update the best, but in some cases it can cause issues). The best practice is to return a large finite value or use a try-except to handle failures. This directly explains the poor best model.

Why this answer

Hyperopt's fmin function minimizes the loss returned by the objective. If the objective returns NaN, Hyperopt may not handle it gracefully; NaN comparisons are always false, so the best loss may not update correctly, and a failed trial could be inadvertently selected. The correct approach is to ensure the objective returns a finite value, such as a large number, when the model fails to train.

This prevents NaN from polluting the search.

Exam trap

The trap here is assuming that Hyperopt automatically discards trials with NaN losses, when in fact NaN can silently break the optimization logic.

27
MCQeasy

A data scientist is training a model on Databricks and wants to track experiments using MLflow. They need to record the model's hyperparameters, evaluation metrics, and the resulting model artifact. They also want to be able to compare runs and reproduce results later. Which MLflow component should they use to organize these runs?

A.MLflow Tracking Server
B.MLflow Projects
C.MLflow Experiment
D.MLflow Model Registry
AnswerC

An MLflow Experiment is the primary organizational unit for runs. It groups related runs, allowing you to track parameters, metrics, and artifacts for each run. By creating an experiment and logging runs within it, you can easily compare results and reproduce experiments. This is exactly what the data scientist needs to organize and manage their model training efforts.

Why this answer

MLflow Experiments are the fundamental unit for organizing runs. Each experiment can contain multiple runs, each capturing parameters, metrics, artifacts, and metadata. This structure enables easy comparison and reproducibility.

While other MLflow components like the Model Registry and Projects play important roles, the experiment is the correct choice for grouping and tracking training runs during model development.

Exam trap

The trap here is confusing the Model Registry with the experiment tracking functionality, as both are part of MLflow but serve different stages of the model lifecycle.

28
MCQeasy

A data science team has deployed a model to Databricks Model Serving and wants to ensure that the endpoint can handle sudden spikes in traffic without manual intervention. Which feature should they configure?

A.A larger workload type with more memory and CPU.
B.Deploy multiple endpoints and use a round-robin DNS.
C.Enable scale-to-zero to reduce cold starts during spikes.
D.Autoscaling with a defined minimum and maximum replica count.
AnswerD

Autoscaling automatically adjusts the number of replicas based on incoming traffic, ensuring the endpoint can handle spikes without manual scaling. By setting minimum and maximum replicas, you bound the scaling to control cost and capacity. This is the intended feature for handling variable load in Databricks Model Serving.

Why this answer

Autoscaling is the Databricks Model Serving feature that dynamically adjusts the number of replicas based on traffic, allowing the endpoint to handle spikes without manual intervention. Configuring minimum and maximum replicas ensures that scaling is bounded and cost-effective. Other options either provide static capacity or are not designed for dynamic load handling.

Exam trap

The trap here is confusing scale-to-zero with autoscaling; scale-to-zero reduces cost during idle periods but can cause cold starts, while autoscaling adds replicas to handle increased load.

29
MCQmedium

Refer to the exhibit. An MLOps engineer is reviewing a JSON object representing a model in the Databricks Model Registry. The engineer wants to promote this model to the 'Production' stage using the MLflow Python API. Which command is correct?

A.client.transition_model_version_stage(name='revenue_forecast', version=5, stage='Production')
B.client.update_model_version(name='revenue_forecast', version=5, new_stage='Production')
C.mlflow.register_model(model_uri='revenue_forecast/5', stage='Production')
D.client.set_model_stage(model='revenue_forecast', ver=5, to='Production')
AnswerA

This method is the correct MLflow API call for changing the stage of a registered model. By specifying the model name, version, and target stage, the engineer programmatically moves the model through the lifecycle, which is a fundamental requirement for automated deployment pipelines in Databricks.

Why this answer

To transition a model version in MLflow, the `transition_model_version_stage` function is the standard method. It requires the model name, the specific version, and the target stage string. This API call is critical for programmatic CI/CD pipelines, allowing engineers to automate the promotion process based on validation results, thereby reducing manual effort and ensuring consistent deployment procedures across the organization.

Exam trap

Candidates often guess the API method name, confusing it with generic MLflow logging methods or incorrectly assuming they need to delete and re-register the model to change its stage.

30
MCQmedium

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with a strict response time SLA of under 50 milliseconds. The model includes an extensive text-cleaning pipeline that utilizes heavy regex matching. How should the engineer package the model to ensure maximum inference efficiency and meet the low-latency requirement?

A.Store the preprocessing logic in a Delta table and have the serving endpoint query the table asynchronously during the scoring request.
B.Develop a separate Azure or AWS Lambda function to handle the text cleaning before forwarding the request payload to the model endpoint.
C.Implement a custom MLflow PyFunc model where both the text preprocessing and the scikit-learn predictor are encapsulated inside the predict method.
D.Deploy the scikit-learn model natively without custom wrapper code and require client applications to execute the regex cleaning logic locally.
AnswerC

Encapsulating preprocessing within a custom MLflow PyFunc guarantees that raw input strings are cleaned and transformed consistently in the same execution context as the model. This eliminates extra network round trips, optimizes memory usage, and ensures predictable sub-50ms inference performance.

Why this answer

Integrating the preprocessing directly into the MLflow PyFunc wrapper ensures that the input transformation runs within the optimized inference container memory space, avoiding out-of-band network calls and reducing serialization overhead. This architectural pattern prevents latency bottlenecks often introduced by separate preprocessing microservices, satisfying strict production SLAs.

Exam trap

Candidates separate preprocessing logic into an external API or upstream service, creating network latency bottlenecks that violate strict response time SLAs.

31
MCQhard

An ML engineer is deploying a model that includes a custom Python class for preprocessing. During deployment to a Model Serving endpoint, the model fails to load with a 'ModuleNotFoundError'. What is the most likely cause of this error despite having the class in the training notebook?

A.The custom class was not saved as a separate .py file and included in the 'code_paths' parameter during logging.
B.The Model Serving endpoint does not support custom Python classes for security reasons.
C.The data scientist forgot to install the 'databricks-model-serving' library in the training cluster.
D.The custom class must be registered in the Unity Catalog as a separate 'Function' entity.
AnswerA

Notebook-defined classes exist only in the memory of the training session. To make them available to the Model Serving endpoint, the class definition must be in a Python file that is explicitly uploaded to the MLflow artifact store along with the model during the logging process.

Why this answer

When MLflow logs a model, it captures the environment but not necessarily the local code or classes defined in the notebook unless they are part of a package or provided as a code dependency. For custom classes to be available in the serving container, they must be included in the 'code_paths' argument of the log_model function.

Exam trap

Candidates assume that because a custom class is defined and works inside an interactive notebook, MLflow will automatically pickle or capture it without explicit file inclusion.

32
MCQhard

When designing a robust MLOps pipeline for high-stakes financial applications, why should you prioritize 'reproducibility' over 'speed' during the deployment phase?

A.Speed is irrelevant in all machine learning contexts.
B.Reproducibility is required for compliance, auditability, and reliable debugging.
C.Databricks does not support fast deployment pipelines.
D.Speed increases the risk of data leakage during training.
AnswerB

Financial regulators require evidence of how model decisions are made. If you cannot reproduce the training process, you cannot guarantee the integrity of the model. Furthermore, debugging production issues is impossible without being able to recreate the exact environment, artifacts, and data state that produced the problematic output.

Why this answer

Reproducibility ensures that a model can be recreated exactly using the same code, environment, and data. In high-stakes fields like finance, being able to audit every decision is a regulatory requirement. While speed is useful, a fast deployment of an un-reproducible model creates significant business risk.

If a model fails, you must be able to recreate the state to identify and fix the bug to remain compliant.

Exam trap

Candidates often choose speed over reproducibility because agile deployment sounds modern, overlooking strict regulatory and auditing requirements inherent in high-stakes financial environments.

33
MCQmedium

A data science team at a retail company has registered a demand forecasting model in the MLflow Model Registry. The model version is currently in the 'Staging' stage and has been validated by the QA team. Before promoting it to 'Production', the ML engineer wants to ensure that the model's performance does not degrade when serving live traffic. They decide to deploy the model to a small percentage of production traffic while continuing to serve the existing model. Which Databricks feature should they use to achieve this?

A.Databricks Model Serving with traffic splitting
B.Databricks Jobs with conditional task execution
C.MLflow Model Registry webhooks
D.MLflow Projects with Docker environments
AnswerA

Databricks Model Serving supports traffic splitting, allowing you to route a percentage of requests to a new model version while the rest go to the existing version. This enables safe canary deployments and A/B testing. By configuring traffic splitting on the endpoint, the team can gradually shift traffic and monitor performance without impacting all users.

Why this answer

Databricks Model Serving natively supports traffic splitting, which lets you distribute incoming requests across multiple model versions on the same endpoint. This is ideal for canary deployments: you can send a small fraction of traffic to the new version, monitor metrics, and then increase the percentage. Other options lack the real-time routing capability required to test a model under live production load without affecting all users.

Exam trap

The trap here is confusing model deployment automation (webhooks, jobs) with traffic management, which requires a serving feature that can split requests between versions.

34
MCQmedium

You maintain a Databricks ML pipeline that trains a model nightly and registers new versions in MLflow Model Registry. A downstream batch scoring job in another workspace loads the model by stage. Auditors require that every production scoring run can be traced back to the exact training data snapshot and code commit. Which approach best satisfies this requirement?

A.Use mlflow.log_input with a Delta table dataset, log the Git commit as a run tag, and register the model version from that run so the run ID links model, data, and code.
B.Store the training data path in the model's signature and use the workspace notebook revision to identify the code commit.
C.Set the model version's description to the Git commit hash and rely on the registered model name to identify the data snapshot.
D.Log the training run with mlflow.log_param for the Git commit and log the Delta table version as a tag, then register the model version from that run.
AnswerA

mlflow.log_input records a dataset entity with its source and digest on the run, and the run ID becomes the immutable link between the model version, the exact data snapshot, and the logged Git commit tag. Because MLflow automatically stores the run ID on the model version, auditors can traverse from a scored model back to the precise training inputs and code revision without relying on human-maintained metadata.

Why this answer

The requirement is an immutable, queryable chain from a production model version to the exact training data snapshot and code revision. Logging the dataset via mlflow.log_input and the Git commit as a run tag, then registering from that run, creates that chain because the model version stores the run ID, and the run stores the dataset digest and code tag. Manual descriptions, parameters, signatures, or notebook revisions are mutable or incomplete and cannot satisfy an audit.

Exam trap

The trap here is assuming that any logged metadata such as a description or parameter creates an auditable link, when only the run ID and logged dataset entities provide an immutable, queryable chain.

35
MCQmedium

You are using MLflow to track experiments for a computer vision model. After several runs, you notice that the training loss is not being logged correctly because the metric name contains a space. What is the recommended way to log metrics with names that include spaces?

A.Use the `mlflow.log_metric` function with the space included; MLflow will automatically sanitize it.
B.Replace spaces with underscores or hyphens before logging.
C.Log the metric as a tag instead of a metric.
D.Use a custom MLflow plugin to allow spaces in metric names.
AnswerB

MLflow metric and parameter names must be valid identifiers; spaces are not allowed. The recommended practice is to sanitize names by replacing spaces with underscores or hyphens. This ensures the metrics are logged correctly and can be queried and visualized without errors in the MLflow UI and API.

Why this answer

MLflow enforces strict naming rules for metrics and parameters, disallowing spaces. The correct approach is to replace spaces with underscores or hyphens before logging. This maintains compatibility with the MLflow tracking server, UI, and API, and avoids logging errors.

Other options either do not solve the problem or introduce unnecessary complexity.

Exam trap

The trap here is thinking that MLflow will automatically handle spaces in metric names or that tags are a suitable substitute for numeric metrics.

36
Multi-Selecthard

When utilizing Hyperopt with MLflow on Databricks for distributed hyperparameter tuning, which TWO components are strictly required to configure the optimization run properly? (Select TWO)

Select 2 answers
A.An objective function that accepts hyperparameter values and returns a loss metric
B.A SparkSession configuration object initialized with custom shuffle partitions
C.A search space definition specifying the hyperparameters and their distributions
D.A pre-registered model URI pointing to an existing MLflow Model Registry entry
E.A delta table containing the engineered features for the validation split
AnswersA, C

The objective function is mandatory: Hyperopt's fmin minimises its returned loss value, which must be derived from the hyperparameters passed in. This drives every trial and determines the best configuration, so it is strictly required alongside the search space.

Why this answer

Hyperopt optimization routines require an objective function that returns a loss value to minimize and a search space definition that dictates the boundaries of the hyperparameters being explored. Mastering these components is essential for conducting scalable, automated machine learning experiments on Databricks clusters.

Exam trap

Candidates often include 'an MLflow experiment' or 'a GPU cluster' as required components, while the core logical requirements for Hyperopt are strictly the objective function and the search space.

37
MCQmedium

You are preparing a feature-engineering job that reads from a Delta table and writes a feature table. The job must run on a schedule and be idempotent so that reruns after a failure do not duplicate data. Which approach should you use?

A.Use a Databricks Job with a notebook that performs a MERGE INTO the feature table using a deterministic key and a timestamp condition.
B.Use a Databricks Job with a notebook that uses DataFrame.write.mode("overwrite") to replace the entire feature table on each run.
C.Use a Databricks Job with a Delta Live Tables pipeline that writes to a streaming table with a defined primary key and applies AUTO CDC.
D.Use a Databricks Job with a notebook that appends new rows and relies on a downstream deduplication step in the serving layer.
AnswerA

MERGE INTO with a deterministic key and a condition ensures that on rerun, existing rows are updated rather than inserted again, making the operation idempotent. This is the standard pattern for scheduled jobs that must tolerate retries and failures without creating duplicates, and it works with Delta Lake's ACID guarantees.

Why this answer

A MERGE INTO operation keyed on a deterministic identifier and a condition ensures that repeated runs update existing records instead of inserting duplicates, satisfying idempotency. This is the recommended pattern for scheduled feature-engineering jobs in Databricks that must be resilient to retries. The other options either rely on downstream fixes, destroy incremental state, or use mechanisms not designed for idempotent batch writes.

Exam trap

The trap here is assuming that any Delta write is automatically idempotent, when in fact append operations will duplicate rows on rerun unless a MERGE or similar upsert logic is used.

38
MCQeasy

A data scientist is using MLflow to track experiments on Databricks. They want to record the value of a hyperparameter named 'learning_rate' for a run. Which MLflow function should they use?

A.mlflow.log_artifact("learning_rate", 0.01)
B.mlflow.log_metric("learning_rate", 0.01)
C.mlflow.set_tag("learning_rate", 0.01)
D.mlflow.log_param("learning_rate", 0.01)
AnswerD

mlflow.log_param() is the correct function to log a single hyperparameter as a key-value pair. It records the parameter under the current run, making it available for comparison in the MLflow UI. This directly satisfies the requirement to record the 'learning_rate' hyperparameter. Other functions log metrics, artifacts, or tags, not parameters.

Why this answer

mlflow.log_param() is specifically designed to log hyperparameters as key-value pairs. It enables filtering and comparison of runs based on parameter values in the MLflow UI. The other functions log metrics, tags, or artifacts, which are not suitable for hyperparameters.

Using the correct function ensures that the 'learning_rate' is properly recorded and can be used for experiment analysis.

Exam trap

The trap here is confusing metrics with parameters, since both accept numeric values, but only parameters are intended for hyperparameters.

39
MCQhard

A machine learning engineer is using MLflow to track experiments on Databricks. They want to ensure that the model's input schema is enforced during inference to prevent errors from malformed data. Which MLflow feature should they use when logging the model?

A.Use `mlflow.register_model()` with a schema definition in the model version description.
B.Log the model with `mlflow.pyfunc.log_model()` and set the `pip_requirements` parameter to include a schema validation library.
C.Log the model with `mlflow.sklearn.log_model()` and set `input_example` to a sample DataFrame.
D.Log the model with `mlflow.pyfunc.log_model()` and provide the `signature` parameter with a `ModelSignature` object.
AnswerD

Providing a `signature` when logging a model with MLflow records the expected input and output schema. During inference, MLflow can validate incoming data against this signature, raising an error if the data does not conform. This helps prevent runtime errors due to schema mismatches and ensures consistent model behavior.

Why this answer

MLflow model signatures define the expected input and output types. When a model is logged with a signature, MLflow can validate input data during inference, ensuring it matches the schema. This prevents errors from malformed data and improves reliability.

The `signature` parameter is the correct way to enforce schema.

Exam trap

The trap here is confusing `input_example` with schema enforcement; `input_example` only helps infer a signature but does not enforce it.

40
MCQmedium

You need to perform cross-validation on a large dataset while ensuring that your model remains performant. What is the Databricks-recommended approach?

A.Collect all data to the driver node and use Scikit-Learn's cross_val_score.
B.Use a Spark-based cross-validator or distribute folds across the cluster.
C.Perform cross-validation only on a small, sampled subset of the data.
D.Hardcode the fold split indices to ensure reproducibility.
AnswerB

Distributing folds across the cluster allows for parallel processing of cross-validation, which is crucial for large-scale Databricks workloads. This approach ensures that the model is thoroughly validated using the full capacity of the cluster, maintaining high efficiency while managing the computational load of training multiple model variations concurrently.

Why this answer

Leveraging Spark's distributed nature for cross-validation allows you to train folds in parallel, maximizing cluster utilization. By using libraries like Spark MLlib or distributed cross-validation wrappers, you can scale the evaluation process to datasets that would otherwise be impossible to handle on a single node. This ensures robust model evaluation while respecting the operational constraints of large-scale distributed data processing.

Exam trap

Candidates often suggest manual loops or single-node Scikit-Learn cross-validation, which fails to utilize the distributed computing power of the Databricks cluster for large-scale dataset processing.

41
Multi-Selectmedium

A team is deploying a scikit-learn model to Databricks Model Serving and wants to minimize cold-start latency so that the first request after a period of inactivity is still fast. Which TWO actions help achieve this? (Choose two.)

Select 2 answers
A.Set the endpoint's scale-to-zero behavior so that instances remain warm and are not fully shut down during idle periods.
B.Reduce the size of the model artifact and its logged dependencies so that container startup and model loading complete faster.
C.Enable Inference Tables on the endpoint so that request payloads are cached and replayed on subsequent cold starts.
D.Increase the number of served model versions on the endpoint so that requests can be spread across more copies of the model.
E.Lower the endpoint's concurrency setting so that each replica handles fewer simultaneous requests and starts faster.
AnswersA, B

Scale-to-zero controls whether the endpoint's compute is torn down when there is no traffic. Disabling it, or configuring the endpoint to keep instances provisioned, means the model stays loaded and the first request after idle time does not pay the container start and model load cost, which is the dominant contributor to cold-start latency.

Why this answer

Cold-start latency comes from provisioning compute and loading the model environment. Keeping instances warm by avoiding full scale-down removes the provisioning cost, and shrinking the model artifact plus its dependencies shortens the load phase. Neither logging nor routing changes affect how quickly a replica becomes ready.

Exam trap

The trap here is confusing Inference Tables, which log payloads for audit, with a caching or warm-up mechanism that reduces cold-start latency.

42
MCQhard

A machine learning engineer is using MLflow Tracking to log metrics and artifacts for a deep learning model. They notice that the training run logs a large number of metrics (e.g., loss per batch) and want to reduce the storage footprint and improve query performance. They also need to retain the ability to compare runs and reproduce results. Which of the following actions is most appropriate?

A.Disable metric logging entirely and rely on TensorBoard for visualization.
B.Log only the final epoch metrics and use MLflow's step parameter to log intermediate metrics less frequently.
C.Use a different tracking server with a NoSQL backend to handle the high volume of metrics.
D.Store metrics in a separate file and log it as an artifact instead of using mlflow.log_metric.
AnswerB

Logging metrics with the step parameter allows you to record a time series, but logging every batch can create a huge number of metric entries. By logging only final epoch metrics or reducing frequency, you decrease storage and speed up queries. MLflow still stores the history if needed, but the volume is manageable. This balances reproducibility with efficiency.

Why this answer

MLflow's log_metric function can accept a step parameter to record metrics over time, but logging every batch creates a large number of entries. Reducing the logging frequency or only logging final epoch metrics significantly decreases storage and improves query speed while still providing enough data for comparison and reproducibility. Other options either lose MLflow's tracking capabilities or introduce unsupported configurations.

Exam trap

The trap here is assuming that any reduction in logged metrics requires abandoning MLflow tracking, when in fact you can simply log less frequently using the step parameter.

43
MCQmedium

A data scientist has completed hyperparameter tuning with Hyperopt on Databricks and now needs to register the best-performing model to the MLflow Model Registry, including its signature and input example, so that downstream scoring jobs can validate incoming data. The training script uses MLflow autologging. Which approach most reliably captures the signature and input example during registration?

A.Use the MLflow Client's update_model_version() method after registration to attach the signature and input example as tags on the model version.
B.Call mlflow.register_model() directly on the run URI returned by the tuning run, because autologging automatically infers and stores the signature and input example.
C.Explicitly call mlflow.log_model() (or the flavor-specific log_model) with the signature and input_example arguments inside the training function, then register that logged model.
D.Register the model from the MLflow experiment's artifact location by copying the MLmodel file and manually editing it to include the signature and input example fields.
AnswerC

Passing signature and input_example to a flavor-specific log_model call records both artifacts with the model version. This is the deterministic way to guarantee the signature schema and a representative input example are attached, which downstream scoring jobs can then use to validate incoming payloads before invoking the model.

Why this answer

The signature and input example are artifacts attached at model logging time, so the reliable path is to pass them explicitly to the flavor-specific log_model call used by the training function. Registering afterward from the run URI, editing tags, or hand-editing the MLmodel file does not produce the complete, tooling-readable artifacts needed for downstream data validation.

Exam trap

The trap here is assuming that MLflow autologging always captures both a signature and an input example, when in practice the input example must be supplied explicitly.

44
MCQmedium

A data scientist is developing a scikit-learn model on Databricks and wants to track the full lineage of the training data, including the exact Delta table version used. They are using MLflow Tracking with a Unity Catalog-enabled workspace. Which approach best captures this lineage as part of the MLflow run?

A.Enable MLflow autologging for scikit-learn, which automatically records the Delta table version used during training.
B.Log the Delta table version as a tag using `mlflow.set_tag` and record the table name in the run's tags.
C.Use `mlflow.log_artifact` to save a copy of the Delta table's transaction log as a JSON file.
D.Rely on Unity Catalog's automatic lineage tracking, which captures data-to-model relationships without additional logging.
AnswerB

Logging the Delta table version and table name as tags directly associates the exact data snapshot with the MLflow run. This enables reproducibility and lineage tracking because anyone can query the run and identify which version of the table was used for training, which is critical for governance and debugging.

Why this answer

The correct approach is to explicitly log the Delta table version and table name as tags on the MLflow run. This creates a clear, queryable link between the model and the exact data snapshot, supporting reproducibility and audit requirements. Neither autologging nor Unity Catalog lineage alone captures the specific version used in a run, and saving transaction logs as artifacts is inefficient and not queryable.

Exam trap

The trap here is assuming that Unity Catalog's automatic lineage or MLflow autologging will record the exact Delta table version without explicit user action.

45
MCQeasy

A data scientist has trained a model and logged it with MLflow. They now need to deploy it as a real-time endpoint on Databricks Model Serving. The model requires a custom Python library that is not available in the default environment. What is the correct way to ensure the library is available when serving the model?

A.Install the library on the driver node of the cluster used for serving, and restart the cluster.
B.Use a custom container image for the model endpoint and install the library at runtime via a startup script.
C.Add the library to the Databricks workspace's global init script so it is available to all clusters.
D.Include the library in the model's conda environment or requirements file when logging the model with MLflow.
AnswerD

When logging a model with MLflow, you can specify a conda environment or a requirements file that lists all dependencies. Databricks Model Serving uses this information to create the serving environment, ensuring the custom library is installed. This is the standard and supported method for managing model dependencies in serving.

Why this answer

The correct way is to include the custom library in the model's conda environment or requirements file when logging with MLflow. Databricks Model Serving uses this specification to build the serving environment, ensuring the library is available. Other methods like cluster-level installations or runtime scripts do not apply to serverless serving endpoints and can introduce reliability issues.

Exam trap

The trap here is assuming that Model Serving endpoints run on clusters you can configure, but they are serverless and dependencies must be declared in the model artifact.

46
MCQmedium

Refer to the exhibit. If the 'train_model' task fails, what happens to the 'evaluate_model' task in this Databricks Job?

A.It executes regardless of the failure.
B.It is skipped because the dependency is not met.
C.It attempts to run using the last successful model artifact.
D.It retries the 'train_model' task automatically until success.
AnswerB

The 'depends_on' field creates a strict prerequisite. When the parent task fails, the DAG execution stops or skips the child task by design. This mechanism protects the pipeline from executing logic based on incomplete or invalid model artifacts, maintaining the reliability of the overall MLOps training workflow.

Why this answer

In Databricks Jobs, a task that depends on another will only execute if the parent task finishes successfully. If 'train_model' fails, the dependency condition for 'evaluate_model' is not met, causing the job to stop or skip the downstream task. This behavior is intentional to prevent evaluating a non-existent or corrupted model, ensuring pipeline integrity and avoiding wasted computational resources on dependent tasks.

Exam trap

Candidates assume the job continues regardless of failure, not realizing that Databricks Jobs default to a strict dependency model where downstream tasks are aborted if the parent fails.

47
MCQmedium

A data science team is transitioning from manual model training to automated pipelines in Databricks. They require a mechanism to track model lineage, versions, and stage transitions programmatically. Which Databricks component best satisfies this requirement?

A.Databricks Jobs Scheduler
B.MLflow Model Registry
C.Databricks Repos
D.Delta Lake
AnswerB

The Model Registry is specifically engineered to handle model versioning, stage transitions (e.g., Staging to Production), and lineage tracking. It provides the necessary APIs to automate these workflows within CI/CD pipelines, ensuring that model deployment follows a robust and governed process in Databricks environments.

Why this answer

MLflow Model Registry provides a centralized model store, set of APIs, and UI for managing the full lifecycle of MLflow models. It enables versioning, model stage transitions, and lineage tracking, which are critical for MLOps maturity. By using the Model Registry, teams ensure reproducibility and governance, allowing for seamless promotion of models from staging to production while maintaining an audit trail of changes and deployment history.

Exam trap

Candidates often confuse MLflow Tracking with MLflow Model Registry. They select Tracking because it records metrics, but it lacks the lifecycle management and stage transition capabilities required for formal model promotion.

48
MCQeasy

When implementing a Feature Store in Databricks, what is the primary benefit of using a Feature Table compared to a standard Delta table for feature engineering?

A.Feature Tables automatically convert all data into Parquet format.
B.They enable automated point-in-time joins for training data generation.
C.They provide faster query performance than standard Delta tables.
D.They allow users to write SQL queries directly to the feature data.
AnswerB

Point-in-time joins are critical for avoiding training-serving skew and data leakage. Feature Tables store metadata that allows the Feature Store to perform these joins automatically, ensuring that the features used for training are temporally consistent and truly represent the state at the time of observation.

Why this answer

Feature Tables provide metadata tracking and lineage, linking features directly to the models that consume them. This metadata enables automated point-in-time joins, preventing data leakage during training by ensuring features are sampled as they existed at a specific timestamp. This level of automation is not available with standard Delta tables, making Feature Tables essential for building reproducible, production-ready machine learning pipelines.

Exam trap

Candidates mistakenly believe that Feature Tables are just a storage format for performance, ignoring the critical architectural capability of point-in-time joins which prevents data leakage during training.

49
MCQmedium

A machine learning engineer wants to ensure that model training artifacts are persistent and accessible even if the ephemeral compute cluster is terminated. What is the standard practice in Databricks for achieving this?

A.Save the model to the local file system using the /dbfs/ path prefix.
B.Log the model using MLflow, which stores the artifact in the configured backend.
C.Export the model artifact to a temporary workspace file and download it locally.
D.Configure the cluster to use an external persistent volume for local storage.
AnswerB

MLflow handles the persistence of model artifacts by automatically uploading them to the managed storage backend (e.g., DBFS or Unity Catalog). This approach ensures that the model is versioned, tracked, and accessible across different clusters or users, maintaining the persistence required for production-grade machine learning lifecycle management.

Why this answer

In Databricks, local cluster storage is ephemeral and is deleted when the cluster terminates. To ensure model artifacts persist, developers must log them to MLflow, which automatically handles the backend storage in DBFS or Unity Catalog-managed storage. This practice decouples the model lifecycle from the compute lifecycle, ensuring that models remain available for deployment or further evaluation after the compute resources are released.

Exam trap

Candidates assume that saving files to the local file system is sufficient. They ignore the fact that cluster storage is ephemeral and will be wiped upon termination.

50
MCQeasy

A data scientist has developed a scikit-learn model and wants to deploy it as a REST API endpoint on Databricks Model Serving. They have logged the model with MLflow and registered it in the MLflow Model Registry. The model requires a specific Python library that is not pre-installed in the Databricks Model Serving environment. What should the data scientist do to ensure the library is available when the model is served?

A.Package the library as a wheel file and upload it to DBFS, then reference it in the serving endpoint configuration.
B.Add the library to the model's conda environment and log the model with that environment.
C.Modify the cluster's Spark configuration to include the library path.
D.Install the library on all cluster nodes using an init script.
AnswerB

Databricks Model Serving uses the conda environment specified in the MLflow model to install dependencies. By including the required library in the conda environment and logging the model with it, the serving environment will install the library automatically. This ensures the model has all necessary dependencies at runtime.

Why this answer

Databricks Model Serving builds the serving environment based on the conda environment recorded with the MLflow model. To include a custom library, you must add it to that conda environment before logging the model. When the model is deployed, the serving infrastructure will create an environment with the specified dependencies.

Other methods like init scripts or Spark config do not apply to the serverless serving environment.

Exam trap

The trap here is assuming that Model Serving uses clusters or can reference external files for dependencies, when it actually relies on the conda environment packaged with the model.

51
MCQmedium

Your company uses MLflow tracking. You want to query all experiments that achieved a specific accuracy threshold across multiple teams. What is the most efficient way to achieve this?

A.Export all experiment metadata to a CSV and use Excel to filter for the desired accuracy.
B.Use the 'mlflow.search_runs()' API to programmatically filter runs based on the accuracy metric.
C.Manually browse through each experiment folder in the MLflow UI to find the metrics.
D.Query the underlying Delta tables in the MLflow experiment storage location directly.
AnswerB

The search_runs API is the standard and most efficient way to query MLflow experiment data. It supports filtering by metrics, parameters, and tags, enabling users to quickly retrieve relevant information programmatically. This method integrates perfectly with automated analysis tools, CI/CD pipelines, and centralized reporting dashboards.

Why this answer

The MLflow Search API (search_runs) is specifically designed to query experiments based on metrics, parameters, and tags. By using the Python MLflow client to filter runs, you can programmatically extract insights across different workspaces. This avoids manual searching, allows for automated report generation, and facilitates governance by allowing you to easily identify top-performing models or flag models that don't meet corporate performance standards.

Exam trap

Candidates often suggest using the UI to manually filter runs, which is inefficient for large-scale enterprise environments compared to the programmatic 'search_runs' API.

52
MCQmedium

You are migrating a legacy ML pipeline to Databricks. You need to ensure that the feature engineering logic used during training is identical to the logic used during real-time inference. What is the recommended approach?

A.Copy the feature transformation code into both the training notebook and the inference service.
B.Use the Databricks Feature Store to define and log features.
C.Store the features in a Parquet file and load it during both training and inference.
D.Implement the transformation logic as a SQL view in the Hive Metastore.
AnswerB

The Feature Store enables the packaging of features with their associated transformation logic. When you use the Feature Store for training, it automatically logs the transformations. During inference, you can then retrieve the features using the same ID, ensuring that the exact same logic is applied to the input data.

Why this answer

Consistency between training and inference (the 'training-serving skew') is a primary cause of model failure. The Databricks Feature Store acts as the single source of truth for feature definitions. By using it, you ensure that the same transformations are applied during both training and inference, eliminating discrepancies that arise from re-implementing logic in different languages or frameworks, which is critical for model reliability.

53
MCQmedium

When utilizing the Databricks Feature Store for model development, why should a developer define a primary key in the Feature Table?

A.To increase the write throughput of the Delta table underlying the feature store.
B.To enable the Feature Store to perform automated lookups for model inference.
C.To bypass the need for performing data validation checks on features.
D.To automatically convert all features into a sparse matrix format.
AnswerB

Primary keys provide the necessary structure for the Feature Store to perform O(1) or O(log n) lookups during online or batch inference. By specifying keys, the system knows exactly how to join features with input observation data, ensuring accurate, automated feature retrieval for real-time model predictions.

Why this answer

Defining a primary key is essential for enabling point-in-time joins during model inference. It allows the Feature Store to perform lookups efficiently and ensures that data consistency is maintained between training and serving. By using primary keys, the system can automatically handle complex join operations, reducing the risk of data leakage and simplifying the feature engineering pipeline for production deployment.

Exam trap

Candidates assume primary keys are only for relational database normalization, overlooking their critical role in automated feature lookups and preventing data leakage.

54
MCQhard

A team is deploying a model to Databricks Model Serving that requires a specific version of a Python library that conflicts with the version pre-installed in the serving environment. They include the library version in the model's requirements.txt. However, upon deployment, the endpoint fails to start, and logs indicate a dependency conflict. What is the most likely cause of this failure?

A.The library version conflicts with a pre-installed library that is required by the serving infrastructure, causing a dependency resolution failure.
B.The model's requirements.txt is not being parsed correctly due to a syntax error.
C.The serving environment ignores requirements.txt and uses only the pre-installed libraries.
D.The specified library version is incompatible with the Python version used by the serving environment.
AnswerA

Databricks Model Serving has a set of pre-installed libraries that are critical for the serving runtime. If your requirements.txt specifies a version that conflicts with these, pip may fail to resolve dependencies, or the endpoint may crash at runtime. This is a common cause of deployment failures when custom dependencies clash with the base environment.

Why this answer

Model Serving environments come with pre-installed libraries that support the serving infrastructure. If a model's requirements.txt specifies a version that conflicts with these, the dependency resolver may fail, preventing the endpoint from starting. The solution is to align the requested version with the pre-installed one or use a custom container image to isolate dependencies.

Exam trap

The trap here is assuming that any library version can be installed, when in fact the serving environment has fixed dependencies that can cause conflicts.

55
MCQmedium

You are designing a strategy for monitoring model performance after deployment. Which of the following is the most important indicator that a model requires retraining?

A.An increase in the number of concurrent inference requests.
B.A significant drop in model prediction accuracy on production data.
C.The expiry of the model's 'Production' stage status.
D.A minor change in the underlying Databricks runtime version.
AnswerB

Model accuracy is the bottom-line metric for performance. A drop in accuracy indicates that the relationship between inputs and outputs has changed, implying that the model is no longer reflecting the current reality of the data. This is the clearest, most urgent signal that the model requires retraining.

Why this answer

Performance degradation, often caused by model drift or data drift, is the primary driver for retraining. While monitoring system metrics like latency is important, the predictive quality of the model is the ultimate metric. Detecting a significant drop in accuracy or precision indicates that the model is no longer meeting business requirements, triggering the need for a new training cycle in an automated MLOps workflow.

56
MCQhard

A machine learning engineer is using MLflow to track experiments on Databricks. They notice that when they run `mlflow.log_artifact` with a local file path inside a notebook, the artifact is stored in the run's artifact location, but when they run the same code in a job cluster, the artifact is missing. The job cluster uses the same MLflow tracking server and experiment. What is the most likely reason for the missing artifact?

A.The MLflow tracking server rejects artifacts from job clusters because they lack an interactive user session, so the artifact is discarded.
B.The local file path used in the job cluster points to a location that is not accessible from the driver, such as a path on the local filesystem of a different node, so the file does not exist when `log_artifact` is called.
C.The artifact location is configured with a short-lived credential that expires before the job cluster completes, causing the artifact upload to fail silently.
D.The job cluster does not have the MLflow library installed, so `log_artifact` silently fails without logging an error.
AnswerB

In a job cluster, code may run on the driver, but if the file is created on an executor's local filesystem or a path that is not synchronized, `log_artifact` on the driver cannot find it. Unlike notebooks attached to an interactive cluster, job clusters do not share local storage across nodes. The file must be written to a distributed location like DBFS or Unity Catalog volume before logging.

Why this answer

In a Databricks job cluster, code may execute on the driver, but local filesystem paths are not shared across nodes. If the artifact file is created on an executor or a different node, the driver cannot access it when `log_artifact` is called. To ensure artifacts are logged, write them to a distributed storage location such as DBFS or a Unity Catalog volume, then log from there.

Exam trap

The trap here is assuming that local filesystem paths behave identically in interactive notebooks and job clusters, overlooking that job clusters do not share local storage across nodes.

57
MCQmedium

You are designing a model retraining strategy. What is the most reliable way to trigger a retraining job based on model performance degradation?

A.Schedule the retraining job to run every Monday regardless of current model performance.
B.Use Databricks Workflows to trigger retraining based on drift detection alerts.
C.Send an email alert to the lead data scientist to manually kick off the job.
D.Write a cron job in the notebook that calculates performance every minute.
AnswerB

This event-driven approach ensures that retraining only occurs when it is objectively needed. By triggering Workflows based on drift detection, you maintain the model's relevance to the current data distribution without wasting compute budget on redundant retraining jobs when the model is still performing adequately.

Why this answer

Linking monitoring metrics to workflow orchestration is the standard MLOps pattern for automated retraining. By configuring a monitor that triggers a Databricks Job when performance drifts below a defined metric (e.g., accuracy), you create a closed-loop system. This ensures that models are updated only when necessary, minimizing cost while maintaining high model performance and reducing manual intervention in the model lifecycle.

Exam trap

Test-takers often assume models retrain automatically based solely on elapsed time, overlooking that event-driven workflows triggered by drift alerts are the best practice.

58
Multi-Selectmedium

When using the Unity Catalog Model Registry, what are the primary advantages of using 'Aliases' over 'Versions' when calling a model from a production application? (Select TWO)

Select 2 answers
A.Aliases allow the application to always point to a stable name like '@prod' instead of a hardcoded version number.
B.Aliases improve performance by caching the model weights on the client side.
C.Aliases allow for easier rollbacks by simply reassigning the alias to a previous version.
D.Aliases are required to enable GPU acceleration for models in Unity Catalog.
E.Only models with an alias can be used in Spark structured streaming jobs.
AnswersA, C

By using a symbolic name, the engineering team can update the model version that the alias points to in the registry. The production application, which calls the alias, will automatically start receiving predictions from the new version without requiring a code redeployment or restart.

Why this answer

Aliases provide a layer of abstraction between the application code and the specific model version. This decoupling is a cornerstone of MLOps, as it allows for seamless updates and rollbacks without modifying the client-side code that consumes the model. It also improves readability by using descriptive names like '@champion'.

Exam trap

Test-takers often recommend hardcoding specific integer version numbers in production applications, leading to brittle codebases that require manual code deployments for every model update.

59
MCQmedium

An ML engineer has deployed a model to Databricks Model Serving and wants to update the endpoint to serve a new model version without changing the endpoint URL or causing downtime. Which approach is correct?

A.Update the existing endpoint's configuration to reference the new model version and apply the update.
B.Use MLflow's `transition_model_version_stage` to move the new version to 'Production', which automatically updates the serving endpoint.
C.Create a new endpoint with the new model version, then delete the old endpoint and update DNS to point to the new endpoint.
D.Modify the model's signature in Unity Catalog to match the new version, and the endpoint will detect the change and update automatically.
AnswerA

Databricks Model Serving allows you to update an existing endpoint's configuration, including the model version, via the REST API or UI. The endpoint URL remains the same, and Databricks performs a rolling update to minimize downtime. This is the standard way to deploy a new model version to an existing endpoint without disrupting service.

Why this answer

To update a serving endpoint to a new model version without downtime, you update the endpoint's configuration to reference the new version and apply the change. Databricks handles the rolling update, ensuring the endpoint URL remains unchanged. Other options would either cause downtime, require client changes, or rely on non-existent automatic updates.

Exam trap

The trap here is believing that model registry stage transitions or signature changes automatically propagate to serving endpoints.

60
MCQeasy

When deploying a model to a Databricks Model Serving endpoint, what is the purpose of the 'Small', 'Medium', and 'Large' workload size settings?

A.They define the maximum number of concurrent requests the endpoint can handle.
B.They determine the geographic region where the model will be hosted.
C.They specify the amount of CPU and memory allocated to each model instance.
D.They select the version of the MLflow library used for deployment.
AnswerC

Each size tier provides a specific amount of RAM and vCPU. A 'Small' instance might be sufficient for a simple linear regression, while a 'Large' instance would be necessary for complex ensembles or models with large memory footprints to ensure they don't run out of memory during execution.

Why this answer

The workload size setting in Databricks Model Serving defines the compute resources (CPU and Memory) allocated to each instance of the model. Choosing the right size is a trade-off between the complexity of the model's computation and the cost of the infrastructure. Larger models or those with heavy preprocessing requirements need more resources to maintain low latency.

Exam trap

Candidates often mistake these settings for 'number of instances' or 'scaling limits'. They are specifically for compute resource allocation (CPU/RAM) per instance to match model complexity.

61
MCQmedium

Refer to the exhibit. A user wants to retrieve the 'accuracy' metric from this run programmatically. Which code snippet correctly accesses this value?

A.mlflow.get_metric('accuracy', run_id='550e8400-e29b-41d4-a716-446655440000')
B.client = mlflow.tracking.MlflowClient(); client.get_run('550e8400-e29b-41d4-a716-446655440000').data.metrics['accuracy']
C.mlflow.search_runs(run_ids=['550e8400-e29b-41d4-a716-446655440000'])['accuracy']
D.mlflow.load_metric('accuracy', '550e8400-e29b-41d4-a716-446655440000')
AnswerB

This method correctly instantiates the tracking client, retrieves the specific run object using its unique identifier, and accesses the nested metrics dictionary. This pattern is essential for developers building custom dashboards or automated model evaluation pipelines that need to extract specific performance data from historical MLflow runs.

Why this answer

The MLflow tracking client provides the `get_run` method, which returns an object containing all run metadata, including metrics. Accessing the dictionary via `data.metrics` is the standard way to retrieve tracked values. This is critical for automated model comparison scripts, where developers must programmatically evaluate multiple runs to select the best-performing iteration for registration in the Model Registry.

Exam trap

Candidates frequently confuse client-side API methods with fluent API calls, or incorrectly attempt to access metrics via dictionary syntax directly on the client object.

62
Multi-Selecthard

Which TWO actions are required to properly implement MLflow Model Registry stages and governance for a machine learning project?

Select 2 answers
A.Assign the 'Can Manage' permission on the Model Registry to all users.
B.Transition model versions between stages like 'Staging' and 'Production'.
C.Apply ACLs to specific models to restrict access and promote actions.
D.Keep all model versions in the 'None' stage for maximum flexibility.
E.Hard-code the model version string into the production application code.
AnswersB, C

Using stages helps manage the lifecycle of a model effectively. It provides a clear signal to downstream applications about which model versions are approved for specific environments, facilitating a clean handoff from development to production and allowing for programmatic automated deployment pipelines.

Why this answer

Implementing Model Registry stages ensures that only validated models move from development to production. By controlling access permissions, organizations enforce a separation of duties, ensuring that data scientists can register models, while only authorized engineers can promote them. This gated workflow is foundational for regulatory compliance and auditability in enterprise MLOps, preventing unauthorized model deployments into downstream production environments.

Exam trap

Candidates frequently select 'registering models' as the governance action, but registration is a development task. Governance specifically requires stage transitions and ACLs to enforce separation of duties.

63
MCQhard

Refer to the exhibit. Which security configuration is the likely culprit for this job failure?

A.Missing cluster-level library permissions.
B.Insufficient Unity Catalog permissions for the job identity.
C.The table is stored in an encrypted S3 bucket.
D.The Delta table has reached its version limit.
AnswerB

The job is executing under a specific identity (service principal). If that identity lacks the necessary 'SELECT' or 'USE' grants in Unity Catalog for the specified schema, access to the underlying Delta table will be blocked, resulting in the reported permission error during the job execution.

Why this answer

This error points to a failure in Unity Catalog access control. In Databricks, when a job attempts to read from a table, it requires specific 'SELECT' permissions on the table and 'USE' permissions on the catalog and schema. If the job service principal is not granted these roles, the query will fail regardless of the workspace-level permissions assigned to the user.

Exam trap

Students frequently blame cluster node sizing or notebook syntax errors, missing the governance layer where job service principals lack Unity Catalog permissions.

64
MCQhard

Refer to the exhibit. What is the purpose of the 'signature' section in this model configuration?

A.To encrypt the model artifacts before they are saved to the registry.
B.To specify the optimization level for the model training process.
C.To ensure that inference requests match the expected data types and structure.
D.To define the hyperparameters used for the model's final evaluation.
AnswerC

The signature acts as a contract between the model and the calling application. By verifying input types, MLflow Model Serving can reject invalid requests early, providing clear error messages. This prevents downstream runtime exceptions in the model inference code, which is essential for stable production-grade model deployment services.

Why this answer

The signature defines the schema of the model's inputs and outputs. This metadata is used by MLflow and model serving infrastructure to validate inference requests before they reach the model. By enforcing types and expected fields, the system prevents runtime errors due to malformed input, which is critical for maintaining robust production inference pipelines in distributed environments.

Exam trap

Candidates often assume the signature is for model performance monitoring or data logging, failing to realize it is a structural contract used for input validation during inference.

65
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default environment and must be installed from a private PyPI repository. Which method ensures the package is available to the model at serving time?

A.Add the package to the cluster's init script and ensure the cluster is attached to the serving endpoint.
B.Include the package as a wheel file in the model's artifact directory and reference it in the model's conda environment.
C.Upload the package to DBFS and use a %pip install command in a notebook before deploying the model.
D.Specify the package's index URL and credentials in the model's conda environment file, and ensure the serving endpoint has network access to the private repository.
AnswerD

Databricks Model Serving allows you to specify a custom index URL and credentials in the conda environment file (e.g., via pip_requirements or conda_env). The serving environment must have network access to the private repository. This method securely installs the package from the private PyPI repository during environment build.

Why this answer

To use a private PyPI repository, the model's conda environment must include the index URL and credentials. Databricks Model Serving builds the environment based on this specification, provided it has network access to the repository. This ensures the custom package is installed and available during inference.

Exam trap

The trap here is assuming that notebook-level installations or cluster init scripts carry over to the serving environment, which is isolated and built from the model's environment specification.

66
MCQmedium

A fraud detection model was deployed to a Databricks Model Serving endpoint. Over the past week, the endpoint's p99 latency has increased from 45 ms to 320 ms, but the model's predictions remain accurate. The serving logs show that each request now includes a larger JSON payload with additional transaction metadata. What is the most likely cause of the degradation?

A.The serving endpoint is hitting the maximum number of concurrent requests and queuing requests, which adds wait time.
B.The model's feature vector has grown, and the scoring function is performing expensive feature transformations on the raw payload before inference.
C.The model artifact is being re-downloaded from the MLflow Model Registry on every request due to a misconfigured artifact cache.
D.The model's underlying compute cluster has been scaled down, reducing available CPU resources for inference.
AnswerB

When the request payload includes extra fields, the model's predict function or preprocessing logic may compute additional features, increasing CPU time per request. This directly explains the latency increase while accuracy stays stable. The other options either would affect accuracy or are not supported by the symptom of larger payloads.

Why this answer

The increased payload size likely triggers additional feature engineering or parsing inside the model's scoring logic, raising per-request CPU time. Since accuracy is unchanged, the model itself is fine; the bottleneck is the preprocessing step. The other explanations either contradict the payload-size correlation or would impact accuracy or availability.

Exam trap

The trap here is assuming that latency degradation always indicates infrastructure scaling problems, ignoring that request payload changes can increase per-request processing time.

67
MCQeasy

When deploying a model as a real-time REST API endpoint on Databricks, which service should be used to manage the serving infrastructure, scaling, and availability?

A.Databricks SQL Warehouse
B.Databricks Model Serving
C.Apache Spark Streaming
D.Databricks File System (DBFS)
AnswerB

Model Serving is the dedicated Databricks product for deploying models as REST APIs. It handles the complexities of scaling, security, and availability, allowing teams to focus on model performance rather than infrastructure management. This is the industry-standard way to expose models within the Databricks ecosystem.

Why this answer

Databricks Model Serving is the managed service designed specifically to deploy models as low-latency, high-availability REST endpoints. It abstracts the underlying infrastructure, providing automatic scaling, version management, and health monitoring. This service is essential for production workloads where reliability and ease of maintenance are paramount, as it removes the burden of manual server administration and infrastructure configuration from the data science team.

Exam trap

Candidates confuse Databricks Model Serving with standard batch jobs or cluster deployments, failing to choose the managed real-time endpoint service.

68
MCQmedium

When deploying a model to a Databricks Model Serving endpoint, how can you ensure the model scales automatically based on traffic demand?

A.Manually adjust the number of instances in the endpoint configuration every time traffic increases.
B.Set the endpoint to a fixed number of instances that is high enough to handle peak traffic.
C.Configure auto-scaling settings in the Model Serving endpoint definition to handle traffic fluctuations.
D.Use a load balancer to redirect traffic to different model endpoints depending on the time of day.
AnswerC

Auto-scaling is a core feature of Databricks Model Serving. By enabling this configuration, the platform automatically manages the allocation of compute resources to match incoming traffic patterns. This provides a balance between high availability during peaks and cost optimization during periods of lower utilization without manual intervention.

Why this answer

Databricks Model Serving utilizes auto-scaling compute resources to handle fluctuations in traffic. By configuring the serving endpoint with appropriate auto-scaling parameters, the platform dynamically adjusts the number of concurrent instances based on request load. This ensures that the application maintains low latency under heavy load while remaining cost-effective during periods of low activity, which is essential for managing production-level model inference at scale.

Exam trap

Candidates often confuse manual endpoint configuration with auto-scaling, mistakenly believing they must manually add nodes or adjust cluster sizes during traffic spikes in the Databricks UI.

69
MCQhard

Refer to the exhibit. Why is the missing model signature considered an MLOps risk in a production environment?

A.It prevents the model from being saved to the local file system.
B.It restricts the model to only be deployed on CPU clusters.
C.It makes it difficult to perform automated input validation at inference time.
D.It slows down the training process by adding extra validation overhead.
AnswerC

Without a signature, the serving system doesn't know the expected input types or shapes. This prevents automated validation, where the system checks data against the schema before passing it to the model. This increases the risk that malformed data will cause the model to crash or produce garbage results.

Why this answer

Model signatures define the expected input schema. Without them, the serving engine cannot validate that incoming data matches what the model expects, which leads to runtime errors or silent, incorrect predictions. In production, this lack of validation makes it impossible to distinguish between data errors and model logic errors, creating significant operational risks and slowing down root cause analysis during incidents.

Exam trap

Candidates often think missing signatures only affect documentation or metadata, failing to realize they are critical for the serving infrastructure to perform automated input validation.

70
MCQmedium

A data scientist has registered a new model version in MLflow Model Registry and wants to validate it against a golden dataset before promoting it. The validation notebook must run automatically whenever a new version is registered in the 'ChurnModel' registry model, and the result should block promotion if the validation fails. Which approach should the data scientist use?

A.Configure a webhook on the MLflow Model Registry that triggers a Databricks job running the validation notebook, and use the job's success/failure to gate promotion.
B.Create a Databricks job with a file arrival trigger that monitors the MLflow artifact store for new model files.
C.Enable automatic model promotion in the registry so that any newly registered version is immediately moved to Production after passing built-in validation.
D.Use MLflow Projects to package the validation notebook and run it manually each time a new version is registered.
AnswerA

MLflow Model Registry webhooks fire on registry events such as MODEL_VERSION_CREATED, allowing an external service (here, a Databricks job) to run validation automatically. The job's outcome can be used to gate promotion, satisfying the requirement for automatic validation upon registration.

Why this answer

MLflow Model Registry webhooks are designed to notify external systems when registry events occur, such as a new model version being created. By configuring a webhook that calls a Databricks job, the validation notebook runs automatically. The job's success or failure can then be used to decide whether to promote the version, providing the required gating mechanism.

Exam trap

The trap here is assuming MLflow Model Registry has built-in automatic validation or promotion, when in fact it relies on webhooks and external orchestration for such workflows.

71
MCQhard

A nightly Databricks job trains a model, registers a new version in Unity Catalog, and then updates a Model Serving endpoint that serves the 'champion' alias. The endpoint must switch to the new version only after the job's validation step passes. Which approach correctly enforces this?

A.Configure the endpoint with a scale-out policy that adds replicas whenever a new model version is registered.
B.Delete the previous model version after training so the endpoint is forced to pick up the newest registered version.
C.Have the job update the endpoint configuration to reference the new model version number directly after training completes.
D.Assign the 'champion' alias to the new version only inside the validation-passed branch of the job, and leave the endpoint configured to serve the alias.
AnswerD

Because the endpoint serves the alias, moving the alias to the validated version is the promotion step, and doing it only after validation passes guarantees the endpoint never serves an unvalidated version. Rollback is a single alias reassignment back to the prior version.

Why this answer

Decoupling promotion from deployment is the key: the endpoint always serves the 'champion' alias, and the job reassigns that alias only after validation succeeds. This makes the alias the single source of truth for what is in production and keeps rollback to a one-step alias change, while the validation branch ensures unvalidated versions never reach the endpoint.

Exam trap

The trap here is thinking that registering a new model version automatically causes a serving endpoint to switch to it, when endpoints only change when their referenced alias or version changes.

72
MCQhard

An ML engineer is training a model with a custom Python loop and wants MLflow to capture training metrics at regular intervals so that partial progress is visible before the run finishes. They are using `mlflow.start_run` and manual logging. Which approach correctly makes intermediate metrics visible during the run?

A.Enable `mlflow.autolog()` and remove the manual logging calls from the loop.
B.Accumulate metrics in a Python list and call `mlflow.log_metrics` once after the loop completes.
C.Call `mlflow.log_metric` inside the loop with the `step` argument set to the iteration number.
D.Set the run tag `mlflow.note.content` inside the loop to the current metric value.
AnswerC

`log_metric` writes the value immediately to the tracking server, and the `step` argument records the iteration index so the metric appears as a time series. Because each call persists right away, dashboards and the run page show progress while the loop is still executing, which is exactly the streaming visibility the engineer wants.

Why this answer

Intermediate visibility comes from persisting each metric as it is produced. Calling `log_metric` inside the loop writes the value immediately and the `step` argument positions it on the metric's time axis, producing a curve that updates live. Deferring logging to a single call at the end, or misusing tags, leaves no partial record if the run is interrupted.

Exam trap

The trap here is assuming autolog or batched logging will surface custom-loop metrics mid-run, when only immediate per-step logging does.

73
MCQhard

A team has a production Databricks Model Serving endpoint for a churn model. They retrain weekly and register new model versions in Unity Catalog. They want the endpoint to automatically pick up the newest registered version without manual intervention, while keeping the previous version available for instant rollback. Which approach should they implement?

A.Enable automatic model version detection on the endpoint so it polls the registry every hour and serves the highest version number.
B.Register each weekly model under the same Unity Catalog model name and rely on MLflow's 'latest' stage to route production traffic automatically.
C.Create a scheduled Databricks job that calls the serving endpoint's update-config REST API with the new version ID whenever a training run completes.
D.Configure the endpoint with the Unity Catalog model name and the alias 'champion'; update the alias to the new version after each weekly training job.
AnswerD

Serving endpoints can reference a Unity Catalog model by name plus an alias, and traffic follows whatever version the alias points to. Updating the champion alias after each training job makes the newest version live without editing the endpoint config, and the prior version stays addressable by its version number or a challenger alias for rollback.

Why this answer

Aliases in Unity Catalog give a stable, named pointer to a model version that a serving endpoint can consume. Promoting a freshly validated version means moving the champion alias to that version, which instantly changes production traffic while leaving earlier versions intact for rollback. This removes manual endpoint edits and custom API automation from the weekly promotion path.

Exam trap

The trap here is assuming an endpoint can auto-discover the newest model version, when serving actually binds to an explicit version or alias that must be promoted.

74
MCQhard

You are monitoring a production model deployed on Databricks Model Serving. You notice that the model's predictions have gradually become less accurate over time, likely due to data drift. You need to implement a solution that automatically detects drift and triggers retraining. Which Databricks feature should you use to monitor the model's input data and performance?

A.Databricks SQL dashboards that query the inference table and display accuracy metrics.
B.MLflow Tracking to log prediction results and manually compare distributions over time.
C.Databricks Lakehouse Monitoring for model quality and data drift, configured on the inference table that logs requests and responses.
D.Unity Catalog lineage to track data transformations and identify changes in input data sources.
AnswerC

Databricks Lakehouse Monitoring can be configured on inference tables to track data drift and model performance metrics over time. It provides automated alerts and can trigger retraining workflows. This is the native Databricks solution for monitoring model quality in production, integrating with the lakehouse for scalable analysis.

Why this answer

Databricks Lakehouse Monitoring is the appropriate feature to monitor model input data and performance for drift. It can be configured on inference tables to automatically compute drift metrics and alert when thresholds are exceeded, enabling timely retraining. Other options lack the automated drift detection and integration with production workflows required for this scenario.

Exam trap

The trap here is confusing data lineage or manual tracking with automated drift detection, but only Lakehouse Monitoring provides out-of-the-box statistical monitoring for production models.

75
Multi-Selecthard

You are implementing a CI/CD pipeline for a Databricks ML project using Databricks Repos and Databricks Asset Bundles. You need to ensure that the pipeline promotes code and model artifacts across dev, staging, and prod workspaces consistently. Which TWO practices should you implement to achieve this? (Choose two.)

Select 2 answers
A.Register model versions to a shared Unity Catalog metastore that is accessible from all workspaces, using environment-specific model names or aliases.
B.Store workspace-specific configuration in `databricks.yml` target blocks and use bundle variables for environment differences.
C.Store model artifacts in a separate cloud storage bucket outside Databricks and copy them manually between environments.
D.Use the same Unity Catalog catalog name in all workspaces and rely on workspace-local paths for artifacts.
E.Hardcode the production workspace URL and cluster ID in the notebook code to simplify deployment.
AnswersA, B

A shared Unity Catalog metastore allows model versions to be registered once and referenced across workspaces. By using environment-specific model names or aliases (e.g., `model_dev`, `model_prod`), you maintain separation while enabling promotion. This aligns with Databricks' recommended governance model and ensures artifacts are consistently available to all environments.

Why this answer

Using Databricks Asset Bundles with target blocks and variables, and registering models to a shared Unity Catalog metastore with environment-specific names or aliases, together provide a consistent, governed promotion path. These practices externalize configuration and centralize artifact management, which are core to reliable CI/CD across multiple workspaces.

Exam trap

The trap here is assuming that using identical catalog names or hardcoding workspace details simplifies multi-environment deployment, when it actually breaks isolation and portability.

Page 1 of 4

Page 2

All pages

Practice Databricks-ML-Pro by domain

Target a specific domain to shore up weak areas.

See all domains with question counts →