Courseiva

Databricks Certified Machine Learning Professional (Databricks-ML-Pro) — Questions 151–225

300 questions total · 4pages · All types, answers revealed

Page 2

Page 3 of 4

Page 4
151
MCQmedium

Your organization requires that all models deployed to production undergo a drift detection check. Which approach is most effective for monitoring model performance in Databricks?

A.Write a custom Spark job to compare the training table with inference logs every hour.
B.Enable Databricks Lakehouse Monitoring on the inference Delta tables to track drift automatically.
C.Configure the Model Registry to automatically retrain the model whenever performance dips below a threshold.
D.Rely on end-user feedback to manually flag when the model performance seems degraded.
AnswerB

Lakehouse Monitoring offers native, integrated drift detection that works directly with Unity Catalog and Delta tables. It provides out-of-the-box dashboards and automated alerts, reducing the need for custom code and ensuring that drift detection is consistently applied across all production models in the environment.

Why this answer

Monitoring drift requires comparing inference data distributions against the training baseline. Databricks Lakehouse Monitoring provides a managed service that automatically detects feature and prediction drift. By leveraging Delta tables as the source of truth, the monitoring service can compute statistics periodically and trigger alerts, ensuring that any degradation in model performance is identified quickly, allowing for proactive retraining or model rollback strategies.

Exam trap

Candidates often suggest building custom monitoring dashboards using SQL or Python code, ignoring that Databricks Lakehouse Monitoring is the native, automated tool specifically built for drift detection.

152
MCQhard

A fraud detection model is deployed to a Databricks Model Serving endpoint. The team wants to test a new model version without affecting existing predictions. They need to send a copy of live traffic to the new version and log its predictions for comparison, while the current version continues to serve all responses. Which feature should they use?

A.Use MLflow's pyfunc flavor to log both models and select the active one at request time.
B.Enable the endpoint's shadow mode by specifying the new model version as a shadow.
C.Deploy the new model version to a separate endpoint and use a load balancer to mirror requests.
D.Configure the endpoint with a traffic split of 50/50 between the current and new model versions.
AnswerB

Databricks Model Serving supports shadow mode, where a copy of the request is sent to a shadow model version while the primary version's response is returned to the client. The shadow model's predictions can be logged for analysis. This exactly matches the requirement to test a new version without affecting existing predictions. The shadow model receives the same input but its output is not served, making it ideal for safe validation.

Why this answer

Shadow mode on Databricks Model Serving allows a new model version to receive a copy of live traffic while the primary version continues to serve all responses. This enables safe evaluation of the new model's predictions without impacting users. Traffic splitting would affect some responses, and external load balancing is not native and adds complexity.

Therefore, enabling shadow mode is the correct approach.

Exam trap

The trap here is confusing shadow deployment with traffic splitting; shadow mode mirrors requests without affecting responses, while traffic splitting changes which model serves the response.

153
MCQmedium

A team is building a feature store in Databricks. They need to ensure that training data and inference data are consistent to avoid training-serving skew. What is the primary benefit of using the Databricks Feature Store in this context?

A.It automatically scales the underlying cluster during training runs.
B.It provides a unified interface for computing features, ensuring consistency between training and inference.
C.It forces data scientists to use only Python for feature engineering.
D.It eliminates the need for any data cleaning or preprocessing steps.
AnswerB

Consistency is achieved by using the same feature definitions and logic for both offline training datasets and online inference lookups. The Feature Store acts as the bridge that guarantees the data fed into the model during training is identical to what the model receives during real-time serving.

Why this answer

The Databricks Feature Store enables the reuse of feature pipelines, ensuring that the exact same transformation logic is applied during both training and inference. By decoupling feature generation from model training, it creates a 'single source of truth' for features, effectively eliminating training-serving skew and improving the reliability of models in production. This is a vital component for maintaining feature consistency across distributed teams.

Exam trap

Candidates often assume the Feature Store is primarily for performance optimization or storage reduction rather than its core purpose of ensuring consistent transformation logic between training and serving.

154
MCQmedium

A machine learning engineer is developing a model on Databricks and wants to ensure that the model's input schema is enforced during inference. They are using MLflow to log the model. What should they do?

A.Log the model with `mlflow.pyfunc.log_model` and provide a `signature` that includes the input schema.
B.Log the model with `mlflow.pyfunc.log_model` and include a custom `predict` method that validates the input schema.
C.Log the model with `mlflow.sklearn.log_model` and set the `input_example` parameter to a sample of the training data.
D.Log the model with `mlflow.sklearn.log_model` and set the `registered_model_name` parameter to register the model in the Model Registry.
AnswerA

Providing a `signature` when logging a model with MLflow defines the expected input and output schema. This signature is used by MLflow to validate input data during inference, ensuring that the data types and column names match. It helps catch schema mismatches early and enforces the contract.

Why this answer

MLflow's model signature defines the expected input and output schema. When a model is logged with a signature, MLflow can validate input data during inference, ensuring that the schema is enforced. This is the standard and most effective way to enforce input schema, as it leverages built-in functionality without custom code.

Exam trap

The trap here is thinking that providing an `input_example` alone enforces schema; it only infers a signature if none is provided, and enforcement requires an explicit signature.

155
MCQmedium

An ML engineer wants to ensure that only models that have passed a specific validation suite can be assigned the 'Champion' alias in Unity Catalog. What is the recommended way to automate this process?

A.Manually checking the validation results and updating the alias in the UI.
B.Using a Databricks Workflow to run validation and then calling the MLflow Client to update the alias.
C.Setting a SQL trigger on the Model Registry table to update the alias automatically.
D.Configuring the Model Serving endpoint to automatically promote the newest version.
AnswerB

A Databricks Workflow can orchestrate the entire process: loading the new model version, running a suite of performance and bias tests, and then using the `set_registered_model_alias` method if the tests succeed. This provides a fully automated, hands-off, and verifiable path to production.

Why this answer

Automation in the model lifecycle is best achieved using Databricks Workflows or CI/CD pipelines. By creating a task that runs a validation notebook, the system can programmatically update model aliases via the MLflow Client API only after all tests pass. This ensures a consistent, governed promotion process that prevents low-quality models from reaching production.

Exam trap

Candidates frequently select manual UI actions or legacy workspace notebooks instead of automated Databricks Workflows combined with the MLflow Client API for governed production promotions.

156
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that uses a custom Python function as a pre-processing step. The function relies on a global variable defined in a separate module. After deployment, the endpoint returns errors indicating the global variable is not defined. The engineer confirmed the module is included in the model's conda environment. What is the most likely cause?

A.Databricks Model Serving runs each inference in a separate process, so global variables are not shared across requests.
B.The model was logged with `mlflow.sklearn.log_model` instead of `mlflow.pyfunc.log_model`, so custom pre-processing code is ignored.
C.The global variable is not serialized with the model because MLflow only saves the model's predict method and its direct dependencies.
D.The conda environment does not include the module because MLflow only captures packages installed via pip, not local modules.
AnswerC

MLflow's default model saving mechanism serializes the model object and its immediate dependencies, but it does not automatically capture global variables or module-level state from custom modules unless they are explicitly referenced within the model's class or function. If the pre-processing function relies on a global variable from another module, that state may not be preserved during serialization, leading to a NameError or undefined variable at serving time.

Why this answer

The error arises because MLflow's serialization does not capture global variables from external modules unless they are explicitly included in the model's code. When the model is loaded for serving, the pre-processing function may reference a global variable that was not saved, resulting in an undefined variable error. To fix this, the engineer should ensure all necessary state is encapsulated within the model or use `code_path` to include the module and initialize variables properly.

Exam trap

The trap here is assuming that any module in the conda environment will have its global state preserved, when in fact MLflow serializes only the model object and its direct code dependencies.

157
MCQhard

You are preparing a model for deployment in a production Databricks environment. Which THREE steps should be included in your model development pipeline to ensure model quality and traceability?

A.Define an input schema using the signature parameter in log_model.
B.Log the model with all temporary files generated during the training process.
C.Use the MLflow Model Registry to promote the model version to Production.
D.Hardcode the model version number directly into the application code.
E.Log all training hyperparameters and evaluation metrics to MLflow.
AnswerA, C, E

Defining an input schema provides metadata that Databricks uses for type validation at inference time. This prevents runtime errors in the serving endpoint when unexpected data types are provided, ensuring the model receives the exact structure it expects for successful prediction calculations and data processing.

Why this answer

A production-ready pipeline requires rigorous validation, tracking, and documentation. Using MLflow to log parameters/metrics ensures reproducibility; setting input schemas allows for automatic validation during serving; and registering models with stage transitions (Staging/Production) ensures that only validated models are exposed to production endpoints, preventing accidental deployment of broken code.

Exam trap

Candidates often omit input schemas or rely solely on manual tracking, forgetting that production readiness strictly requires automated validation signatures and registry governance.

158
MCQmedium

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model's inference function requires access to an external feature store table for real-time feature lookup. Which approach allows the model to retrieve these features during serving while maintaining low latency and avoiding per-request authentication complexity?

A.Package the feature store data as a static file within the model artifact and load it during initialization.
B.Use the Databricks Feature Store's online table and call the feature lookup client within the model's predict method.
C.Enable the model to query the Delta table directly using Spark during each inference request.
D.Configure the serving endpoint to call an external REST API that returns feature values, using a personal access token stored in the environment.
AnswerB

Databricks Feature Store provides an online table that is optimized for low-latency lookups. The feature lookup client can be embedded in the model's predict method, and the serving endpoint automatically authenticates to the online store. This approach avoids managing credentials per request and ensures the model uses consistent feature values.

Why this answer

The Databricks Feature Store online table is purpose-built for low-latency feature serving. By embedding the feature lookup client in the model's predict method, the endpoint can retrieve features in real time without managing authentication. This integration is a core pattern for deploying models that rely on feature store data, ensuring consistency between training and serving.

Exam trap

The trap here is assuming that any external data source can be queried directly from the model without considering the serving environment's limitations and the need for optimized, authenticated access.

159
MCQmedium

A data scientist is using MLflow on Databricks to tune a scikit-learn GradientBoostingRegressor with Hyperopt. They configure fmin with max_evals=50, but notice that runs appear in the experiment without parameters or metrics logged, and the best model cannot be reproduced. They want to ensure every trial is fully tracked. Which change should they make?

A.Use `SparkTrials` instead of `Trials` so that each trial is logged as a separate MLflow run.
B.Call `mlflow.autolog()` before Hyperopt and set `nested=True` in fmin.
C.Wrap the objective function's training and evaluation code in `with mlflow.start_run():` (or use `mlflow.start_run(nested=True)` inside the objective) and log params/metrics explicitly.
D.Increase `max_evals` to at least 200 so that MLflow has enough data to create runs automatically.
AnswerC

Hyperopt's fmin does not automatically create an MLflow run for each trial. Without an active run context, calls to mlflow.log_param or mlflow.log_metric attach to the parent run or fail silently, leaving trials untracked. Starting a run inside the objective ensures each trial gets its own run with parameters and metrics, enabling reproducibility and comparison across trials.

Why this answer

Hyperopt's fmin does not manage MLflow runs; each trial must explicitly start a run to log parameters and metrics. Wrapping the objective in mlflow.start_run ensures isolated tracking per trial, which is essential for reproducibility and comparison. Other options confuse search breadth, autologging, or distributed execution with run lifecycle management.

Exam trap

The trap here is assuming that Hyperopt automatically creates MLflow runs for each trial, when in fact the objective function must manage the run context.

160
MCQeasy

A data scientist wants to use MLflow to track a scikit-learn model training run on Databricks. They call `mlflow.sklearn.autolog()` before training. Which of the following will MLflow automatically log for this run?

A.The source code of the training script and the Git commit hash.
B.The model signature, input examples, and the trained model artifact.
C.The hyperparameter tuning search space and the best parameters from a grid search.
D.The feature importance plot and confusion matrix for classification models.
AnswerB

`mlflow.sklearn.autolog()` automatically logs the model signature, input examples, and the trained model artifact, along with parameters and metrics. This reduces manual logging and ensures reproducibility. The signature and input examples are inferred from the training data. This is the primary benefit of autologging for scikit-learn models.

Why this answer

`mlflow.sklearn.autolog()` automatically logs parameters, metrics, the model signature, input examples, and the trained model artifact. It does not log source code, Git commit hash, hyperparameter search spaces, or custom plots like feature importance. This automation streamlines experiment tracking and ensures that key model metadata is captured without manual intervention.

Exam trap

The trap here is assuming that autologging captures everything about the training process, including source code and custom visualizations, when it actually focuses on parameters, metrics, and the model artifact.

161
MCQhard

A company uses Databricks Model Serving to host a real-time model. They need to perform A/B testing between two model versions, sending 10% of traffic to a new version and 90% to the current version. Which approach should they use?

A.Log both models as a single MLflow model that internally routes based on a random number.
B.Use MLflow Model Registry stages to mark one version as Production and the other as Staging, then rely on the serving endpoint to automatically split traffic.
C.Deploy both model versions to the same endpoint and use the traffic split feature in Model Serving.
D.Create two separate endpoints, one for each version, and use an external load balancer to split traffic.
AnswerC

Databricks Model Serving supports serving multiple model versions on a single endpoint with configurable traffic percentages. This allows seamless A/B testing by routing a percentage of requests to each version. The feature is built into the endpoint configuration, enabling easy adjustments and monitoring.

Why this answer

Databricks Model Serving allows multiple model versions on one endpoint with specified traffic percentages, making A/B testing straightforward. This native feature handles routing without external components. Other approaches require additional infrastructure or custom logic, which are unnecessary and less efficient.

Exam trap

The trap here is assuming that MLflow stages automatically control traffic splitting, when in fact traffic split must be explicitly configured on the serving endpoint.

162
MCQmedium

Your team uses Databricks Model Serving to host a production model. You need to ensure zero-downtime updates while maintaining the ability to revert to the previous version instantly if performance degrades. Which deployment strategy should you implement?

A.Directly update the existing model serving endpoint to the new version.
B.Use a Blue-Green deployment strategy with traffic splitting.
C.Delete the current serving endpoint and create a new one with the updated model.
D.Increase the instance count of the existing endpoint before updating.
AnswerB

Blue-Green deployment enables side-by-side versions, allowing for controlled traffic shifting. This methodology provides a seamless transition path and an instantaneous rollback capability, which are essential requirements for maintaining service reliability and managing risk in high-stakes production machine learning environments.

Why this answer

Blue-green deployment allows you to maintain a production endpoint while simultaneously deploying a new version. By using weighted traffic splitting, you can shift traffic gradually and monitor performance metrics. If the new version exhibits anomalies, you can instantly revert traffic to the stable version, ensuring high availability and minimizing the impact of potential production failures, which is critical for robust MLOps workflows.

Exam trap

Candidates often confuse 'Blue-Green' with 'A/B testing'. While they share traffic splitting, Blue-Green specifically focuses on seamless cutovers and instant rollbacks for deployment reliability, not just statistical comparison.

163
MCQhard

A financial services company uses Databricks Model Serving to deploy a real-time fraud detection model. The endpoint is configured with scale-to-zero enabled. During a period of no traffic, the endpoint scales down to zero. When a sudden burst of requests arrives, the first few requests experience high latency. Which mechanism is responsible for this behavior?

A.The requests are being queued because the endpoint's maximum concurrency is set too low.
B.The endpoint is experiencing a cold start as it provisions resources and loads the model into memory.
C.The model's dependencies are being downloaded from PyPI on each request.
D.The model's container image is being rebuilt from scratch for each request.
AnswerB

When scale-to-zero is enabled, the endpoint reduces to zero replicas during idle periods. Upon receiving new requests, it must provision new instances, pull the container image, and load the model. This cold start introduces latency for the initial requests until the endpoint is fully warmed up.

Why this answer

Scale-to-zero reduces cost by scaling down to zero replicas when idle. When traffic resumes, the endpoint must cold start: provision compute, pull the image, and load the model. This causes latency for the first few requests.

The other options describe issues that would persist regardless of scaling, such as concurrency limits or per-request builds.

Exam trap

The trap here is confusing cold start latency with concurrency or image rebuild issues, which are unrelated to scaling from zero.

164
MCQmedium

A team is transitioning from manual experimentation to a formal MLOps pipeline. Which component should be prioritized to ensure that data used during training is reproducible?

A.Store all training data as CSV files in DBFS with unique filenames for each run.
B.Use Delta Lake time travel to query the training data at the specific timestamp of the model training run.
C.Ask the data engineering team to copy the data to a private folder for every run.
D.Only rely on the model performance metrics to infer the quality of the training data.
AnswerB

Delta Lake time travel allows for precise data versioning. By referencing a specific timestamp or version, the pipeline ensures that the exact training set used in a previous experiment can be perfectly reproduced, which is a fundamental requirement for reliable and compliant MLOps workflows.

Why this answer

Delta Lake's time-travel capability is the standard for ensuring data reproducibility in MLOps. By using 'VERSION AS OF' or 'TIMESTAMP AS OF', you can query the exact state of the training dataset at any point in time. This is critical for auditing, model retraining, and debugging, as it eliminates the ambiguity of changes made to the underlying data sources over time.

Exam trap

Candidates often assume saving a standard CSV file or taking a static snapshot is sufficient, forgetting about Delta Lake's native time-travel features.

165
MCQeasy

A data scientist has a model registered in Unity Catalog and wants to let an external application score it over HTTPS without embedding Databricks credentials in the application. The application's identity is already a service principal in the workspace. Which approach should be used to authenticate calls to the Model Serving endpoint?

A.Configure the endpoint to allow unauthenticated access and rely on network isolation to prevent unauthorized callers.
B.Generate a Databricks personal access token for a workspace user and hardcode it in the external application's configuration file.
C.Embed the workspace URL and a Unity Catalog metastore credential in the request body so the endpoint can verify the caller's identity from the payload.
D.Have the external application authenticate as the service principal and obtain an OAuth token from Databricks to present as a bearer credential on each request.
AnswerD

Databricks supports OAuth machine-to-machine authentication, where a service principal exchanges its client ID and secret for a short-lived OAuth token that is sent as a bearer credential. This gives the external application its own non-human identity with rotatable secrets and avoids embedding a user's personal credentials.

Why this answer

Databricks Model Serving endpoints accept OAuth bearer tokens, and service principals can perform machine-to-machine OAuth to obtain them. That gives the external application a distinct, auditable identity with rotatable secrets, avoiding a hardcoded personal token or an unauthenticated endpoint.

Exam trap

The trap here is treating network isolation or a credential placed in the request body as authentication, when the endpoint validates a bearer token from the Authorization header.

166
MCQeasy

A data scientist is training a model with scikit-learn on Databricks and wants to track the experiment using MLflow. They call mlflow.start_run() and then train the model. After training, they call mlflow.log_param() and mlflow.log_metric(), but later find that the run is not visible in the MLflow experiment UI. What is the most likely reason?

A.They forgot to call mlflow.end_run() to close the run, so it remains in a running state and is not displayed.
B.They did not set the MLflow tracking URI or experiment, so the run was logged to a different experiment or the default location.
C.They called mlflow.log_param() and mlflow.log_metric() outside of the active run context, so the data was discarded.
D.The model training did not produce any metrics, so MLflow skipped creating the run.
AnswerB

If the tracking URI or experiment is not explicitly set, MLflow logs to the default experiment (often ID 0) or to a local file store. On Databricks, the default is usually the workspace's default experiment, but if the user changed the experiment or is using a different tracking server, the run may be in another experiment. This is the most common cause of runs not appearing where expected.

Why this answer

The most common reason for not seeing a run in the expected experiment is that the tracking URI or experiment was not set correctly. MLflow defaults to a specific experiment, and if the user is looking at a different one, the run appears missing. Explicitly setting the experiment with mlflow.set_experiment() or the tracking URI ensures the run is logged to the intended location.

Other options would typically cause errors or still show the run.

Exam trap

The trap here is assuming that forgetting to end the run hides it, when in fact running runs are visible; the real issue is often misconfigured experiment or tracking URI.

167
MCQmedium

Your team is migrating models to Databricks Model Registry. You need to ensure that production models are only deployed after passing specific validation tests. Which feature best facilitates this automated governance?

A.Using MLflow Experiments to log model artifacts.
B.Hard-coding model versions in inference notebooks.
C.Configuring Model Registry webhooks for status transitions.
D.Executing manual SQL queries against the MLflow backend.
AnswerC

Webhooks allow you to register external endpoints that are triggered when specific events occur, such as a model version transition to 'Production'. This enables automated CI/CD integration, ensuring that only verified models move forward, which is a foundational requirement for robust, automated MLOps governance workflows.

Why this answer

Model Registry webhooks are the critical component for automating governance. By triggering external CI/CD pipelines (like Jenkins or GitHub Actions) upon model version transitions (e.g., 'Pending' to 'Production'), organizations enforce quality gates. This ensures that only models validated by automated unit tests or performance benchmarks are promoted, reducing manual intervention and minimizing the risk of deploying underperforming models into production environments.

Exam trap

Candidates often assume manual review processes or standard notification emails can enforce automated governance, ignoring the role of Model Registry webhooks in triggering CI/CD pipelines.

168
MCQmedium

A data scientist needs to perform batch inference on a large dataset stored in Delta Lake using a model registered in the Unity Catalog. Which approach is most efficient for leveraging Spark's distributed computing capabilities while using the MLflow model?

A.Loading the model with mlflow.pyfunc.load_model and using a for-loop to iterate over DataFrame rows.
B.Using mlflow.pyfunc.spark_udf to wrap the model and applying it to the Spark DataFrame columns.
C.Converting the Delta table to a Pandas DataFrame and using the standard model.predict() method.
D.Calling the Model Serving REST API for every row in the Delta Lake table using a standard Python request.
AnswerB

The spark_udf function automatically handles the distribution of the model environment and weights to all worker nodes. It allows the model to process data partitions in parallel, significantly reducing the time required for batch inference on massive Delta Lake tables while maintaining a simple, high-level API.

Why this answer

MLflow provides a built-in function to load models as Spark User Defined Functions (UDFs). This allows the model to be distributed across the executor nodes of a Spark cluster, enabling parallel processing of large datasets. This method is preferred over manual iteration as it integrates seamlessly with the Spark DataFrame API and optimizes resource utilization during high-volume batch jobs.

Exam trap

Candidates often try to use standard Python loops or UDFs that do not leverage Spark's parallelization, which is inefficient for large datasets and fails to utilize cluster resources.

169
MCQeasy

When working in Databricks, where should a data scientist primarily look to monitor the resource utilization and execution logs of an active model training job?

A.The MLflow Model Registry.
B.The Databricks Jobs UI.
C.The Databricks Filesystem (DBFS) root directory.
D.The Workspace 'Notebook' sidebar.
AnswerB

The Jobs UI provides a comprehensive view of execution logs, status, and cluster metrics for scheduled or triggered training jobs. It is the central place to monitor the health and performance of the infrastructure during a training run, making it the correct tool for troubleshooting resource-related issues.

Why this answer

The 'Jobs' UI and the Spark UI are the primary interfaces for monitoring training jobs. While MLflow provides tracking for metrics and parameters, the Jobs UI provides the necessary system-level visibility into cluster performance, execution status, and logs. This is critical for troubleshooting memory errors, diagnosing bottlenecks, and managing costs during resource-intensive training sessions.

Exam trap

Candidates confuse MLflow's UI with the cluster management UI. They look for resource utilization metrics in MLflow, which tracks model performance, not the underlying compute cluster's hardware health.

170
Multi-Selectmedium

Which TWO of the following are primary benefits of using the MLflow Model Registry in Databricks?

Select 2 answers
A.Automated training of models.
B.Lifecycle state tracking (e.g., Staging, Production).
C.Centralized versioning and lineage.
D.Real-time monitoring of inference latency.
E.Automatic data cleaning for training sets.
AnswersB, C

Tracking the state of a model version is a core feature of the registry. It allows teams to clearly define which models are ready for production, which are being tested, and which have been deprecated, ensuring that deployment pipelines only pull models from the appropriate, authorized lifecycle stages.

Why this answer

The Model Registry centralizes model management, providing a single source of truth for model lifecycle stages and version history. This is vital for MLOps, as it ensures that teams can track which models are in production, who approved them, and how they perform over time, enabling consistent deployment workflows and improved auditability for regulatory compliance and internal governance.

Exam trap

Candidates sometimes select 'model training' or 'artifact storage' as primary benefits, confusing general MLflow features with the specific governance and lifecycle management provided by the Model Registry.

171
MCQmedium

A fraud detection team has a model registered in Unity Catalog as main.ml.fraud_model. They need to serve it in real time, but compliance requires that every scoring request automatically generate an audit record in a Delta table, and that the model only be promoted to production after a human reviews the audit logs from a canary period. Which deployment configuration satisfies these requirements?

A.Deploy the model with Databricks Model Serving and configure the endpoint to write request logs to a mounted cloud storage path using a cluster-scoped log4j appender.
B.Deploy the model with Databricks Model Serving and enable Inference Tables, then use the endpoint's request/response logs to drive the review before promoting the model version via a UC alias.
C.Deploy the model behind a Databricks SQL warehouse by registering it as a Python UDF, and rely on the warehouse's query history as the audit trail.
D.Deploy the model with Databricks Model Serving and enable model monitoring, which persists every raw request and response to a Delta table for human review.
AnswerB

Inference Tables capture the payloads and responses of every scored request into a Delta table automatically, which is exactly the audit record compliance needs. Because the endpoint references a UC model version through an alias, the team can point the production alias at the canary version only after reviewing those logs, giving a clean, reversible promotion path.

Why this answer

Inference Tables are the Databricks mechanism that automatically logs the request and response payloads of a served model into a Delta table, which directly produces the compliance audit trail. Pairing that with a Unity Catalog alias lets the team hold production traffic on the prior version and repoint the alias only after reviewing the canary logs.

Exam trap

The trap here is assuming that enabling model monitoring also captures raw request and response payloads, when monitoring only derives metrics from the Inference Table that must be enabled separately.

172
MCQhard

A machine learning engineer is using MLflow to log a model trained with a custom algorithm. They want to ensure that the model can be served with a specific input schema and that the schema is enforced during inference. Which MLflow feature should they use?

A.MLflow run tags
B.Model registry stage transitions
C.Model signature with input and output schema
D.MLflow projects with a conda environment
AnswerC

A model signature defines the expected input and output schema, including column names and data types. When logging a model with a signature, MLflow validates that the input at inference time matches the schema. This enforces the contract and prevents errors due to mismatched data. It is the standard way to ensure schema enforcement.

Why this answer

The model signature is the MLflow feature that captures the expected input and output schema. When a model with a signature is served, MLflow validates incoming data against the schema, rejecting mismatches. This ensures that the model receives data in the correct format and types.

Other features like tags, stages, and projects serve different purposes in the lifecycle.

Exam trap

The trap here is confusing model packaging and lifecycle features with schema enforcement, when only the model signature provides input validation at inference time.

173
MCQmedium

A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune an XGBoost classifier. After several trials, they notice that each trial runs on a single executor and the overall tuning job takes much longer than expected. They want to speed up hyperparameter tuning without changing the search space. Which adjustment is most likely to improve performance?

A.Enable autoscaling on the cluster and set `parallelism` to 1 in `SparkTrials`.
B.Increase the `max_evals` parameter in the `fmin` call to run more trials.
C.Increase the number of Spark executors and set `parallelism` in `SparkTrials` to a value greater than 1.
D.Switch from `SparkTrials` to `Trials` and run the search on the driver node only.
AnswerC

SparkTrials distributes trials across Spark executors when parallelism is greater than one. By increasing executors and setting parallelism, multiple hyperparameter configurations run concurrently, reducing wall-clock time. This is the intended use of SparkTrials for single-node ML libraries like XGBoost, where each trial is independent. The search space remains unchanged, so model quality is not compromised.

Why this answer

SparkTrials parallelizes hyperparameter trials across Spark executors when `parallelism` is set above 1. The data scientist observed that each trial used a single executor, indicating that parallelism was effectively 1 or the cluster lacked executors. By adding executors and increasing parallelism, multiple trials run simultaneously, cutting total tuning time.

The search space and model logic remain unchanged, so this is the direct fix.

Exam trap

The trap here is assuming that merely adding cluster resources will speed up SparkTrials without adjusting the parallelism parameter that governs concurrent trials.

174
MCQhard

An ML engineer is training a PyTorch model on a Databricks cluster and wants to automatically log training metrics, parameters, and the model artifact to MLflow without writing explicit mlflow.log_* calls in the training script. The engineer also needs the run to be nested under a parent run that tracks the overall experiment. Which approach should the engineer use?

A.Use mlflow.pytorch.autolog() and rely on it to automatically create a nested run for each training session.
B.Manually log all metrics and parameters using mlflow.log_metric() and mlflow.log_param(), and use mlflow.start_run(nested=True) for nesting.
C.Use the MLflow Tracking API to create a parent run, then call mlflow.pytorch.autolog() inside the parent run without starting a child run.
D.Enable autologging with mlflow.pytorch.autolog() before training, and use mlflow.start_run(nested=True) to create a child run under the parent run.
AnswerD

mlflow.pytorch.autolog() automatically captures metrics, parameters, and model artifacts during training without manual logging calls. Wrapping the training in mlflow.start_run(nested=True) creates a child run nested under an active parent run, satisfying the nesting requirement. This combination directly addresses both needs: automatic logging and hierarchical run organization.

Why this answer

Autologging for PyTorch is enabled via mlflow.pytorch.autolog(), which captures training metrics, parameters, and the model artifact automatically. To nest the run under a parent, the engineer must explicitly start a child run with mlflow.start_run(nested=True). Combining these two features satisfies both automatic logging and hierarchical run organization without manual logging calls.

Exam trap

The trap here is assuming that autologging automatically creates nested runs or that manual logging is required for nesting.

175
MCQeasy

A team notices that a deployed Model Serving endpoint occasionally returns stale predictions for a subset of customers shortly after the nightly feature refresh completes. The offline feature table is confirmed current. Which is the most likely explanation?

A.The offline feature table's Delta transaction log has not been checkpointed, so readers see an older snapshot.
B.The online feature store was not republished after the nightly refresh, so lookups return the previous feature values for those customers.
C.The model version serving traffic is older than the latest registered version, so predictions lag behind the current model.
D.The endpoint's autoscaling added replicas that cached outdated feature values during scale-out.
AnswerB

The online store is a materialized copy that only reflects the offline table when a publish runs. If the nightly job updates the offline table but no publish follows, the endpoint continues serving the older values. Republishing after the refresh aligns the online store with the current offline data and eliminates the stale predictions.

Why this answer

Feature Store keeps an offline table for training and an online table for serving. Updates to the offline table do not propagate automatically; a publish step materializes them into the online store. When the nightly refresh runs without a subsequent publish, the endpoint keeps reading the prior values, producing stale predictions for affected customers until the online store is refreshed.

Exam trap

The trap here is assuming the online feature store stays synchronized with the offline table without an explicit publish step.

176
MCQmedium

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with maximum throughput and minimum latency. The model requires an external preprocessing Python script during inference. Which deployment approach best leverages MLflow and Databricks architecture?

A.Log the model and preprocessing code using custom pyfunc, register it in Unity Catalog, and configure a serverless Model Serving endpoint.
B.Deploy the raw scikit-learn model artifact to Model Serving and execute the preprocessing logic inside a separate Spark structured streaming job.
C.Write the preprocessing code inside a Databricks notebook and call the model endpoint using a client-side REST API request.
D.Save the model weights to DBFS and build an external Flask application on Amazon EC2 to manage traffic routing and inference.
AnswerA

Custom pyfunc models encapsulate both the estimator and arbitrary preprocessing code into the artifact. Registering this artifact in Unity Catalog enables secure governance, and deploying it to a serverless endpoint provides auto-scaling and low-latency inference capabilities.

Why this answer

Packaging the custom preprocessing logic along with the scikit-learn estimator into an MLflow PyFunc model ensures that all inference data transformations happen inside the serving container natively. This prevents client-side processing bottlenecks and guarantees consistent feature engineering between training and real-time serving environments.

Exam trap

Candidates often suggest performing preprocessing in the Spark application before calling the model. This creates training-serving skew and latency issues compared to native PyFunc model packaging.

177
MCQeasy

A data scientist is using MLflow to track a training run. They want to log a dictionary of hyperparameters and a list of evaluation metrics that are computed at the end of each epoch. Which MLflow API calls should they use to log these items?

A.Use `mlflow.log_param()` for each hyperparameter individually and `mlflow.log_metric()` for each metric, calling the latter once per epoch with the epoch number as the step.
B.Use `mlflow.log_artifact()` to log a JSON file containing the hyperparameters and metrics, and then parse it in the MLflow UI.
C.Use `mlflow.log_params()` for the hyperparameters and `mlflow.log_metrics()` for the metrics, calling the latter once per epoch with the epoch number as the step.
D.Use `mlflow.set_tag()` to log the hyperparameters and `mlflow.log_metric()` for the metrics, calling the latter once per epoch with the epoch number as the step.
AnswerC

`mlflow.log_params()` accepts a dictionary of parameters and logs them as key-value pairs. `mlflow.log_metrics()` accepts a dictionary of metric names to values and an optional `step` argument to record metrics at different points, such as epochs. This is the standard way to log per-epoch metrics for time-series visualization in the MLflow UI.

Why this answer

The correct approach is to use `mlflow.log_params()` for the hyperparameter dictionary and `mlflow.log_metrics()` for the metrics dictionary, specifying the epoch as the step. This leverages MLflow's batch logging APIs and ensures that metrics are recorded as time-series data, which the UI can plot over steps. It also keeps parameters organized in the parameters section.

Exam trap

The trap here is thinking that logging hyperparameters as tags or artifacts is sufficient, when in fact MLflow treats parameters, metrics, and tags separately, and only parameters and metrics are comparable in the UI's main views.

178
MCQmedium

When logging a model that requires custom libraries (e.g., a specific version of a non-standard package), how do you ensure the environment is reproducible on the serving endpoint?

A.Install the dependencies manually on the serving cluster after deployment.
B.Include the library binaries directly in the model artifact folder.
C.Provide an explicit conda.yaml or requirements file during the log_model call.
D.Rely on the serving cluster's default environment libraries.
AnswerC

Providing an explicit environment file during the log_model call ensures that MLflow captures the exact dependency requirements for the model. This allows the serving infrastructure to build a matching environment automatically, guaranteeing that the model runs in a configuration identical to the one used during training and validation.

Why this answer

Including a conda.yaml file or an MLflow requirements file ensures that the exact environment dependencies are captured alongside the model. When the model is deployed to a serving endpoint, Databricks uses this metadata to recreate the identical software stack. This prevents the 'it works on my machine' problem, ensuring that inference logic is executed in a consistent, predictable environment across all deployment stages.

Exam trap

Candidates rely solely on the default environment captured by the notebook kernel, forgetting that serving endpoints require explicit dependency files like conda.yaml.

179
Multi-Selectmedium

When developing a model, which THREE actions should a data scientist perform to ensure the model is ready for production deployment via Model Serving?

Select 3 answers
A.Log the model with a defined input/output signature.
B.Record all training parameters and metrics using MLflow.
C.Register the model to the Unity Catalog Model Registry.
D.Hardcode the data path to an external S3 bucket within the model object.
E.Include the entire training dataset in the model's metadata.
AnswersA, B, C

The signature provides a schema for the model, which is essential for request validation in Model Serving. Without a signature, the inference service cannot verify incoming data, leading to cryptic errors or silent failures during batch or real-time scoring when incompatible data formats are provided by upstream systems.

Why this answer

Production-ready models require strict standards: a defined schema (signature) for input validation, proper metadata logging for reproducibility, and registration in the Model Registry. These steps ensure the model is governable, traceable, and secure. Skipping any of these components makes the model difficult to debug, deploy reliably, or monitor for performance drift in production environments.

Exam trap

Candidates often mistake model logging for simple code saving. They forget that production-readiness requires specific schema validation (signature) and centralized governance via the Unity Catalog, rather than just local artifact storage.

180
MCQmedium

Which statement best describes the role of the 'Signature' in an MLflow model when deploying to a Databricks Model Serving endpoint?

A.It is a cryptographic hash used to verify the integrity of the model files.
B.It specifies the data types and names of the input features and the output predictions.
C.It lists the specific users and groups who have permission to call the endpoint.
D.It contains the hyperparameters used during the final training run of the model version.
AnswerB

The signature serves as a contract between the model and the client. It ensures that the serving endpoint can reject malformed requests (e.g., a string where a float is expected) before they cause a crash in the model's prediction code, improving the overall robustness of the service.

Why this answer

An MLflow Model Signature defines the schema of the inputs and outputs for a model. This is crucial for Model Serving as it allows the endpoint to perform input validation before the data reaches the model. It also enables the UI to generate a 'Call' template, making it easier for developers to test the API.

Exam trap

Candidates often confuse MLflow Signatures with model artifacts or performance metrics, believing signatures dictate model accuracy or handle training data transformations automatically during inference, rather than focusing solely on schema definition.

181
MCQeasy

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. What is the simplest way to create the endpoint?

A.Create a Databricks job that runs the model's predict method on a schedule.
B.Write a Python script using the MLflow library to deploy the model to a local server.
C.Use the Databricks UI to create a new serving endpoint and select the registered model.
D.Export the model as a pickle file and upload it to a Databricks cluster.
AnswerC

The Databricks UI provides a straightforward, guided workflow to create a serving endpoint by selecting a registered model from Unity Catalog. It automatically configures the endpoint with the model's environment and dependencies, making it the simplest method for deployment.

Why this answer

The Databricks UI offers a simple, integrated way to create a Model Serving endpoint by selecting a registered Unity Catalog model. It handles configuration and deployment automatically, making it the easiest method for data scientists.

Exam trap

The trap here is confusing model deployment methods, such as local serving or batch jobs, with the managed Model Serving endpoint creation process.

182
Multi-Selectmedium

Which THREE features are provided by the Databricks Model Registry for model lifecycle management?

Select 3 answers
A.Transitioning models between stages (e.g., Staging to Production).
B.Automatic retraining of models on a fixed daily schedule.
C.Tracking version history of registered models.
D.Automated deployment of models to edge devices like mobile phones.
E.Commenting and collaboration on specific model versions.
AnswersA, C, E

Stage transitions are central to the CI/CD model deployment workflow. They allow organizations to promote models through a well-defined lifecycle, ensuring that models in 'Production' have passed all necessary testing, validation, and approval steps, which is a requirement for enterprise-grade deployment and risk management in machine learning.

Why this answer

The Model Registry provides the governance layer for machine learning, allowing teams to track stages, version history, and approval workflows. These features are necessary to transition from an experimental notebook-based workflow to a production environment where changes to models are managed, tested, and audited, ensuring that only validated artifacts are promoted to critical business-facing applications.

Exam trap

Test-takers often select hyperparameter tuning or data preprocessing options, confusing operational tracking features with the Model Registry's lifecycle governance tools.

183
MCQmedium

Which action should be taken to ensure that sensitive PII data is not leaked when logging models using MLflow?

A.Encrypt the entire MLflow experiment folder using workspace-level encryption.
B.Ensure data used for model artifacts is scrubbed of PII.
C.Set the MLflow model registry to 'Private' mode.
D.Delete the MLflow run immediately after deployment.
AnswerB

The most effective way to prevent leakage is to ensure that no PII is included in the training sets or metadata captured by MLflow. Data engineering pipelines should perform masking or removal before the data reaches the training notebook, effectively ensuring the model artifacts are compliant with privacy standards.

Why this answer

Sanitizing data before logging or excluding sensitive columns from the training set is a security best practice. MLflow logs metadata and often sample data. If PII is included, it is persisted in the artifact store.

By ensuring only non-sensitive features are logged, you comply with data privacy regulations and prevent unauthorized access to sensitive information within your model artifacts.

Exam trap

Candidates often suggest masking data at the serving layer or using encryption, forgetting that MLflow logs the training artifacts, so PII must be scrubbed before the logging process begins.

184
MCQhard

When deploying a model to a production endpoint, your team requires that the model is served from a specific, immutable version. How do you ensure this in MLflow?

A.Always point the endpoint to the 'Production' tag.
B.Specify the exact model version in the serving configuration.
C.Use the latest model artifact from the experiment folder.
D.Rebuild the model container on every request.
AnswerB

Hard-coding the version number in the serving config ensures the endpoint is immutable. Regardless of what happens to the 'Production' alias in the registry, the endpoint will continue to serve that specific version, which is the required approach for high-stakes production environments requiring strict version control.

Why this answer

Linking the deployment to a specific Model Registry version alias or version number is the only way to ensure immutability. When you serve a model, you should point to a specific version (e.g., version 5) rather than a dynamic alias like 'Production' if you need to guarantee that no subsequent transitions will change the behavior of the endpoint, thereby ensuring production stability.

Exam trap

Candidates often rely on dynamic tags like 'Production' alias, missing the requirement for strict immutability which demands referencing an exact version number.

185
MCQeasy

Which component of MLflow tracks the parameters, metrics, and tags associated with a specific training run?

A.MLflow Model Registry.
B.MLflow Tracking.
C.MLflow Projects.
D.MLflow Recipes.
AnswerB

MLflow Tracking provides the API and UI to log and visualize parameters, code versions, metrics, and output files. It acts as the central repository for experiment metadata, enabling teams to track and compare the performance of various model iterations effectively throughout the development cycle.

Why this answer

MLflow Tracking is the primary API used for logging and querying machine learning experiments. It allows data scientists to record key information during the training phase, which is vital for experiment reproducibility and comparing performance across different model iterations. Without MLflow Tracking, maintaining visibility into the history of experiments would be manual, fragmented, and prone to losing critical context.

Exam trap

Candidates often confuse MLflow Tracking with the MLflow Model Registry, incorrectly assuming that tracking is responsible for stage management or deployment rather than just recording experiment metadata.

186
MCQmedium

When hyperparameter tuning using `mlflow.spark.autolog()` or `hyperopt`, what is the primary advantage of logging the parameters to the MLflow tracking server?

A.It automatically triggers a new training run when the parameters are updated.
B.It allows developers to visualize and compare results across hundreds of tuning iterations.
C.It compresses the model weights to reduce the storage footprint on DBFS.
D.It prevents overfitting by forcing the model to use regularization parameters.
AnswerB

The primary benefit is the ability to query and visualize the relationship between hyperparameters and metrics. Using the MLflow UI, developers can sort, filter, and plot these parameters, which is essential for identifying the best-performing model from a large space of candidates generated during hyperparameter optimization.

Why this answer

Logging parameters allows for comprehensive lineage and comparison of model runs. It enables the 'experimentation' phase to be evidence-based, where developers can correlate parameter configurations with model performance metrics. This is crucial for reproducibility and for identifying the optimal configuration that maximizes model performance, as it creates a permanent, queryable history of the entire tuning process.

Exam trap

Test-takers often assume autologging directly improves model accuracy, rather than recognizing its true value in providing traceability and comparative analysis.

187
MCQeasy

You need to ensure that a model deployed to a Databricks Model Serving endpoint can be rolled back quickly if it starts performing poorly. Which feature should you use?

A.Use MLflow Model Registry to transition the model version back to Staging, which automatically reverts the serving endpoint.
B.Delete the current model version from the registry and register the previous version as a new version.
C.Configure the endpoint to use a webhook that triggers a rollback job when performance metrics drop.
D.Enable multiple model versions on the endpoint and use traffic splitting to shift traffic back to a previous version.
AnswerD

Databricks Model Serving supports serving multiple model versions on a single endpoint and allows you to configure traffic splits. If the new version performs poorly, you can quickly shift traffic back to the previous version without redeploying, enabling fast rollback. This is the recommended approach for safe deployments and rollbacks.

Why this answer

Model Serving endpoints can host multiple model versions and use traffic splitting to route requests. If a new version underperforms, you can instantly shift traffic back to a previous version by adjusting the split, achieving a fast rollback without redeployment. Other methods are manual, slow, or do not directly control the serving endpoint.

Exam trap

The trap here is thinking that changing a model version's stage in the registry automatically updates the serving endpoint, but stage transitions and endpoint configurations are separate.

188
MCQeasy

A machine learning engineer needs to schedule a Databricks job that runs a Python wheel task to execute a packaged training pipeline on a recurring basis. The pipeline code is built into a wheel and stored in Unity Catalog volumes. Which job configuration correctly executes this workload?

A.Create a job with a notebook task that runs %pip install on the wheel and then calls the pipeline via subprocess from within the notebook.
B.Create a job with a Spark submit task that passes the wheel as a --py-files argument and runs the training script as the main class.
C.Create a job with a notebook task that contains a single cell calling dbutils.library.install on the wheel path, then imports the training module.
D.Create a job with a Python wheel task, specifying the wheel file location and the entry point as package_name.module_name, and attach a cluster or serverless compute.
AnswerD

A Python wheel task is designed exactly for this: you provide the wheel path and a fully qualified entry point such as package.module, and Databricks installs the wheel on the compute before invoking the entry point. This yields reproducible, versioned execution of the packaged training pipeline on a recurring schedule.

Why this answer

A Python wheel task is the purpose-built job task type for running packaged Python code: you point it at the wheel and a fully qualified entry point, and Databricks handles installation and invocation. Notebook-based installs or subprocess calls, and Spark submit tasks, do not provide the same reproducible, entry-point-driven execution.

Exam trap

The trap here is treating any method that installs a wheel as equivalent to a Python wheel task, when only the wheel task uses the packaged entry point directly.

189
Multi-Selectmedium

A machine learning team is using MLflow on Databricks to manage experiments. They want to ensure that their model training runs are reproducible and that they can compare different runs effectively. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Set the experiment name to the current date to avoid confusion.
B.Store the training dataset in the MLflow run's artifact repository for every run.
C.Use a single MLflow run for all experiments to simplify tracking.
D.Log the evaluation metrics for each run using `mlflow.log_metrics()`.
E.Log all hyperparameters used in each run using `mlflow.log_params()`.
AnswersD, E

Logging metrics allows quantitative comparison across runs. mlflow.log_metrics() records metrics such as accuracy, loss, or AUC at each step or epoch. This data is used in the MLflow UI to visualize performance, select the best run, and detect overfitting. Without metrics, you cannot objectively evaluate or compare models, making it a fundamental practice for experiment tracking.

Why this answer

Logging hyperparameters and evaluation metrics are fundamental for reproducibility and comparison. Hyperparameters record the configuration, while metrics provide quantitative performance. Together, they enable filtering, sorting, and selecting the best runs.

Other practices like using a single run, storing full datasets, or date-based naming do not support effective experiment tracking.

Exam trap

The trap here is overlooking that logging both parameters and metrics is necessary; one without the other limits reproducibility and comparison.

190
MCQeasy

A machine learning engineer is using MLflow to log a model. They want to include custom preprocessing logic that is not part of the model's native library. Which MLflow model flavor should they use to package the model with custom code?

A.`mlflow.sklearn`
B.`mlflow.lightgbm`
C.`mlflow.pyfunc`
D.`mlflow.tensorflow`
AnswerC

The `mlflow.pyfunc` flavor allows you to create a custom Python function model that can include arbitrary preprocessing, postprocessing, and inference logic. You define a class that inherits from `mlflow.pyfunc.PythonModel` and implement the `predict` method. This is the correct choice for packaging custom code with the model.

Why this answer

The `mlflow.pyfunc` flavor is designed for custom Python models. It allows you to define a class with a `predict` method that can include any preprocessing, inference, and postprocessing steps. This makes it ideal for packaging models with custom logic that is not supported by native flavors.

Exam trap

The trap here is assuming that native flavors like sklearn or tensorflow can accommodate arbitrary custom code; they are limited to their respective libraries.

191
MCQhard

A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs of a scikit-learn model and automatically log the best model to the Model Registry. They use `mlflow.sklearn.autolog()` and then call `mlflow.sklearn.log_model` with `registered_model_name`. However, they notice that the model version in the registry does not include the signature or input example. Which action should they take to ensure the signature and input example are logged?

A.Register the model without a signature, then manually update the model version metadata in the Model Registry UI to add the signature and input example.
B.Set the environment variable `MLFLOW_LOG_MODEL_SIGNATURE` to `true` before training, which enables automatic signature logging for all models.
C.Use `mlflow.models.infer_signature` on the training data and then call `mlflow.sklearn.log_model` with the inferred signature, but input example cannot be logged for scikit-learn.
D.Pass a `signature` and `input_example` to `mlflow.sklearn.log_model` explicitly, because autolog does not capture them for scikit-learn models.
AnswerD

While `mlflow.sklearn.autolog()` logs parameters, metrics, and the model, it does not automatically infer and log a model signature or input example for scikit-learn models. To include these, the data scientist must explicitly pass `signature` and `input_example` arguments to `log_model`. This ensures the registered model version has the necessary metadata for schema validation and deployment.

Why this answer

Autolog for scikit-learn logs parameters, metrics, and the model artifact but does not automatically capture a model signature or input example. To include these, the engineer must explicitly pass `signature` and `input_example` to `mlflow.sklearn.log_model`. This is essential for model serving and validation, as the signature defines the expected input schema and the input example provides a sample for testing.

Exam trap

The trap here is believing that autolog captures all metadata including signature, when in fact signature and input example must be explicitly provided for scikit-learn models.

192
MCQmedium

A data science team uses MLflow Tracking with a remote tracking server backed by a Databricks-hosted MySQL instance for the backend store and an Azure Data Lake Storage Gen2 path for artifacts. A model-training notebook writes metrics and a model artifact, then calls mlflow.register_model to promote the run into Unity Catalog. Reviewers report that the run's metrics appear in the experiment UI, but the model version in the registry cannot be loaded by the deployment job. The deployment job fails when it tries to download the artifact. Which action most directly resolves the deployment failure?

A.Re-run the training notebook with mlflow.set_tracking_uri pointing at a local file path so the artifact is written to the driver's disk.
B.Grant the deployment job's service principal READ privileges on the ADLS Gen2 artifact location, or configure a storage credential and external location that the deployment job can access.
C.Register the model again using the run ID instead of the artifact URI so the registry stores the artifact inline in the backend database.
D.Increase the --model-serve-timeout on the deployment cluster so the job has more time to fetch the artifact from the tracking server.
AnswerB

Artifacts are stored at the artifact location configured on the tracking server, not inside the backend database. When the deployment job lacks permission to that ADLS path, it cannot download the model, even though metadata and metrics are visible. Aligning the service principal's access or wiring a Unity Catalog storage credential and external location resolves the download failure.

Why this answer

MLflow separates metadata from artifacts: the backend store holds run and metric records, while the artifact store holds the model files. A deployment job that can read experiment metadata still needs its own credentials to the artifact location. Aligning the service principal's access, or exposing the path through Unity Catalog storage credentials and external locations, is what lets the artifact download succeed.

Exam trap

The trap here is assuming that because metrics and run metadata are visible in the UI, the artifacts are equally accessible to every consumer of the run.

193
Multi-Selecthard

A regulated financial services firm must prove that every model promoted to production on Databricks is traceable and governed. They use Unity Catalog for models and MLflow for experiment tracking. Which two practices most directly satisfy an auditor's requirement to trace a production model version back to its training data and code? (Choose two.)

Select 2 answers
A.Grant the production service principal MANAGE on the registered model so it can update versions as needed.
B.Enable the MLflow run's source notebook to be stored with the run so the exact code revision is retrievable from the run details.
C.Log the training dataset path or Delta table version and the Git commit SHA as tags and parameters on the MLflow run that produced the registered model version.
D.Increase the model version's stage description with a free-text summary of the business purpose and the approving manager's name.
E.Configure the serving endpoint to log every request and response payload to a Delta table for later inspection.
AnswersB, C

MLflow records the source notebook or Git reference with each run, which lets reviewers retrieve the code that generated the model. Combined with the run-to-version link in the registry, this closes the loop from production model back to the authored code, which is precisely what a traceability audit expects.

Why this answer

Traceability requires machine-readable links from the production model version back to the exact data and code. Recording the dataset version and Git commit on the source run, plus preserving the run's source notebook, gives auditors a verifiable chain. Descriptions, request logging, and permissions document or control access but do not establish provenance, so they cannot satisfy the requirement alone.

Exam trap

The trap here is treating documentation fields like version descriptions or access permissions as equivalent to recorded data and code lineage.

194
MCQmedium

A data scientist is building a model that requires custom preprocessing logic that is not available in standard libraries. They need to ensure this logic is bundled with the model for inference. What is the recommended approach to encapsulate this custom logic?

A.Save the preprocessing logic in a separate Python script.
B.Implement a custom class inheriting from mlflow.pyfunc.PythonModel.
C.Use a Spark UDF for preprocessing during inference.
D.Hardcode the preprocessing logic directly into the SQL query.
AnswerB

Inheriting from PythonModel allows developers to define a custom 'predict' method. This method can include any necessary preprocessing or post-processing logic, ensuring that the model acts as a self-contained unit that receives raw data and produces the final output without requiring external code dependencies.

Why this answer

Using the MLflow PyFunc (Python Function) flavor is the standard way to package arbitrary logic with a model. It allows developers to define a custom wrapper that includes preprocessing, prediction, and post-processing steps. When the model is logged, the custom class and its environment dependencies are saved, ensuring that the exact same logic is executed during inference, regardless of the deployment target.

Exam trap

Candidates often mistakenly select generic deployment options or standard framework saving methods, forgetting that custom preprocessing logic requires the specialized MLflow PyFunc model flavor wrapper.

195
MCQhard

An ML engineer is instrumenting a Databricks job that trains and evaluates several models. They want each model's metrics, parameters, and artifacts grouped so that a downstream automated promotion step can compare candidates within the same experiment. Which MLflow practice best supports programmatic comparison across runs in this job?

A.Log all candidates as versions of a single registered model and compare them by reading each version's description field for metric values.
B.Write metrics to a Delta table outside MLflow and have the promotion step read that table, using MLflow only to store the model artifacts.
C.Create a separate experiment for each model candidate so that runs are isolated and cannot interfere with one another.
D.Log all candidates into a single experiment, tagging each run consistently so the promotion step can filter and rank runs by metric within that experiment.
AnswerD

Keeping candidates in one experiment lets the promotion step search runs by metric and filter with tags in a single query, producing a clean ranking of candidates. Consistent tags add the metadata needed to identify job, dataset, or model family, so automated comparison and selection are straightforward and auditable.

Why this answer

Grouping candidate runs in one experiment lets MLflow's search API rank and filter them by metric and tag in a single query, which is what an automated promotion step needs. Splitting experiments, overloading registry descriptions, or duplicating metrics elsewhere all break that integrated, queryable comparison.

Exam trap

The trap here is assuming that MLflow model versions carry metric data that can be ranked, when metrics live on runs within an experiment.

196
MCQeasy

A data scientist has trained a model and wants to register it in the MLflow Model Registry on Databricks. They want to indicate that the model is ready for testing in a pre-production environment. Which stage should they transition the model version to?

A.Archived
B.Staging
C.None
D.Production
AnswerB

Staging is designed for models that are being tested or validated before production. It serves as a pre-production environment where the model can be evaluated with real or simulated traffic. Transitioning to Staging aligns with the goal of testing the model in a controlled setting before promoting it to Production. This is the correct stage for pre-production testing.

Why this answer

The MLflow Model Registry provides stages to manage model lifecycle. Staging is specifically meant for models that are undergoing testing or validation before production. By transitioning to Staging, the data scientist signals that the model is ready for pre-production evaluation, allowing stakeholders to test it without affecting live traffic.

Exam trap

The trap here is confusing the Staging stage with Production or None, but Staging is the correct stage for pre-production testing.

197
MCQeasy

A machine learning engineer is training a model using MLflow on Databricks and wants to compare multiple runs to select the best hyperparameters. They need to view metrics across runs in a single interface. Which MLflow feature should they use?

A.MLflow Projects
B.MLflow Model Registry
C.MLflow Tracking UI
D.MLflow Recipes
AnswerC

The MLflow Tracking UI provides a visual interface to compare runs, including metrics, parameters, and artifacts. It allows sorting and filtering runs by metrics, making it easy to identify the best hyperparameters. This is the standard tool for experiment comparison in MLflow. It directly addresses the need to view and compare multiple runs.

Why this answer

The MLflow Tracking UI is designed to display and compare runs, including metrics, parameters, and artifacts. It allows sorting and filtering, which is essential for selecting the best hyperparameters. Other MLflow components like Model Registry, Projects, and Recipes serve different purposes and do not provide a comparative run view.

Thus, the Tracking UI is the correct tool.

Exam trap

The trap here is confusing the Model Registry with the Tracking UI; the Registry manages model versions, not experiment run comparisons.

198
MCQhard

An ML engineer is training a model on Databricks using MLflow and wants to ensure that the training process is deterministic across runs. They set the random seed for NumPy, Python, and the machine learning framework. However, they observe that the model's performance varies slightly between runs on the same data and cluster configuration. Which factor is most likely causing the non-determinism?

A.Non-deterministic operations in the machine learning framework, such as GPU-accelerated training with cuDNN, which may not be fully deterministic even with seeds set.
B.The cluster's autoscaling feature causing different numbers of executors for each run.
C.The use of `mlflow.autolog()` which introduces randomness in logging.
D.The MLflow tracking server not being configured with a persistent backend store.
AnswerA

Many deep learning frameworks use cuDNN for GPU acceleration, which can be non-deterministic by default due to atomic operations and algorithms that are not reproducible. Even with seeds set, cuDNN may choose different algorithms for convolution or other operations, leading to slight variations. To enforce determinism, you must set framework-specific flags (e.g., torch.use_deterministic_algorithms(True)) and possibly disable cuDNN benchmarking.

Why this answer

Non-determinism in deep learning on GPUs often arises from cuDNN's non-deterministic algorithms. Even with seeds set, operations like convolutions may produce slightly different results due to atomic operations or algorithm selection. To achieve determinism, you must set framework-specific flags to enforce deterministic algorithms and disable benchmarking.

Other factors like autologging or tracking server configuration do not affect model training randomness.

Exam trap

The trap here is assuming that setting random seeds alone guarantees determinism, while GPU-accelerated operations may still introduce non-determinism.

199
MCQhard

You are managing a Databricks environment and need to ensure that ML models are reproducible across different workspaces. Which strategy is most effective for cross-workspace model promotion?

A.Download the model artifact from the source workspace and upload it to the target workspace UI.
B.Use a central MLflow Model Registry in Unity Catalog to share model versions across workspaces.
C.Export the model as a pickled Python object and email it to the operations team.
D.Re-run the training notebook in each workspace to recreate the model artifact locally.
AnswerB

Unity Catalog's centralized model registry provides a single source of truth for all models, regardless of which workspace they are accessed from. This allows teams to promote models through environments (dev, staging, prod) while maintaining full auditability, lineage, and consistent governance over the model's lifecycle.

Why this answer

Using a centralized MLflow Model Registry via Unity Catalog allows models to be shared across workspaces securely. This avoids the manual export/import of artifacts and ensures that lineage, tags, and versioning remain consistent. By managing access through Unity Catalog permissions, you ensure that only authorized environments can read or promote specific models, creating a unified, compliant, and highly scalable model deployment lifecycle.

Exam trap

Candidates often suggest manual artifact copying or exporting/importing pickle files, failing to realize that Unity Catalog provides a centralized, secure, and native way to share models across workspaces.

200
MCQeasy

A data scientist wants to automate the retraining of a model whenever new data arrives in a Delta table. They need to orchestrate a multi-step workflow that includes data validation, feature engineering, model training, and deployment. Which Databricks feature should they use?

A.Databricks Jobs with a multi-task workflow
B.Databricks Repos
C.MLflow Projects
D.Delta Live Tables
AnswerA

Databricks Jobs support multi-task workflows, allowing you to define a directed acyclic graph of tasks with dependencies. You can schedule the job to trigger on a file arrival event or a schedule, and each task can run a notebook or Python script. This is the native orchestration tool in Databricks for building and automating ML pipelines, including retraining and deployment steps.

Why this answer

Databricks Jobs with multi-task workflows is the built-in orchestration service that allows you to define dependencies, schedule triggers, and monitor execution of complex ML pipelines. It can be triggered by file arrival events, making it ideal for retraining when new data lands. Other options like MLflow Projects, Delta Live Tables, and Repos serve different purposes and lack native orchestration capabilities.

Exam trap

The trap here is confusing code packaging or data pipeline tools with orchestration; only Databricks Jobs provides the scheduling and dependency management needed for end-to-end ML workflows.

201
MCQmedium

Your team uses MLflow Model Registry. A model version currently in Production has a critical flaw and must be rolled back to a previous version. The previous version is in the Archived stage. What is the most operationally sound approach to restore service quickly while preserving the audit trail?

A.Create a new registered model with the previous version's artifacts and point the serving endpoint to it.
B.Transition the flawed Production version to Archived, then transition the previous version from Archived to Production.
C.Delete the flawed Production model version and re-register the previous version as a new model version.
D.Leave the flawed version in Production and instead update the serving endpoint to load the previous version by its run ID.
AnswerB

Archiving the flawed version removes it from active serving while retaining its metadata and lineage, and moving the prior version back to Production restores the known-good artifact. This preserves a complete audit trail of stage transitions, which is essential for governance and post-incident review.

Why this answer

The correct approach is to archive the problematic Production version and promote the prior known-good version back to Production. This restores service using an already-validated artifact while maintaining full stage-transition history for audit and rollback traceability. Deleting versions, creating new models, or bypassing the registry all sacrifice governance or introduce unnecessary operational risk.

Exam trap

The trap here is assuming that a rollback requires creating a new model version or deleting the bad one, when stage transitions alone can restore service while preserving lineage.

202
MCQeasy

You are monitoring a model served on a Databricks Model Serving endpoint. You need to track the distribution of incoming request payloads to detect data drift. Which Databricks feature should you use to automatically capture and store inference logs for analysis?

A.Enable inference tables on the model serving endpoint.
B.Configure the endpoint to log to MLflow Tracking.
C.Use Databricks SQL dashboards to query the endpoint's access logs.
D.Enable model serving logs to be written to a cloud storage bucket.
AnswerA

Inference tables in Databricks Model Serving automatically capture the request payloads, response payloads, and metadata for each request to the endpoint. These logs are stored in a Delta table that you can query for monitoring and drift detection. This feature is designed specifically for capturing inference data without additional code.

Why this answer

Inference tables are a Databricks Model Serving feature that automatically logs request and response payloads to a Delta table. This enables easy querying and monitoring for data drift without writing additional code. Other options either do not capture payloads automatically or are not designed for this purpose.

Exam trap

The trap here is confusing MLflow Tracking with inference logging; MLflow Tracking is for training runs, not for capturing live serving requests.

203
Multi-Selecthard

A team is using Databricks Feature Store to manage features for a real-time fraud detection model. They need to ensure that the features used during training are consistent with those served at inference time. Which two actions should they take to achieve this? (Choose two.)

Select 2 answers
A.Publish the model with the Feature Store, so that the model automatically looks up the latest feature values from the online store at inference time.
B.Manually copy the feature values from the offline store to the online store before each inference request to ensure freshness.
C.Use the FeatureStoreClient to create a training set that joins features from the feature tables, ensuring point-in-time correctness.
D.Implement a custom UDF in the model to compute features on the fly from raw data, bypassing the Feature Store entirely.
E.Train the model using a separate notebook that reads directly from the online store to simulate inference-time feature retrieval.
AnswersA, C

When you log a model with Feature Store, it records the feature lookups. At inference time, the model uses the online store to fetch the latest feature values, ensuring that the same feature transformations are applied. This maintains consistency between training and serving and is the recommended practice for real-time models.

Why this answer

Using FeatureStoreClient.create_training_set ensures point-in-time correctness during training, and publishing the model with Feature Store ensures that inference uses the same feature lookups from the online store. Together, these actions guarantee that features are consistent between training and serving, which is critical for real-time fraud detection.

Exam trap

The trap here is thinking that manual synchronization or custom UDFs can replace the Feature Store's automated consistency mechanisms, when they actually introduce skew.

204
MCQmedium

You are developing an MLflow project and want to ensure that your code is reusable. What is the benefit of defining an MLproject file?

A.It automatically generates documentation for the training code.
B.It prevents the model from being registered in the Model Registry.
C.It packages code and dependencies for reliable, reproducible execution.
D.It allows the code to run directly on the Databricks SQL Warehouse.
AnswerC

The MLproject file serves as a manifest that captures the environment and entry points for a project. By defining clear dependencies, it ensures that anyone who runs the code gets the same results, which is foundational for collaborative model development in Databricks where consistency across team members is paramount.

Why this answer

An MLproject file defines the entry points, dependencies, and environment for an MLflow project, effectively packaging the code as a self-contained unit. This is essential for reproducibility, as it allows others to execute the project with a single command, automatically setting up the environment and executing the code exactly as the author intended, regardless of the individual's specific machine environment.

Exam trap

Candidates often confuse the MLproject file with a simple requirements.txt or a notebook, missing that the project file is specifically designed to define entry points and environment reproducibility.

205
MCQeasy

Which component of Databricks is specifically designed to prevent training-serving skew by ensuring that feature engineering code is consistent during both model training and real-time inference?

A.MLflow Tracking Server
B.Databricks Feature Store
C.Unity Catalog
D.Databricks SQL
AnswerB

The Feature Store provides a unified API to compute and store features. When a model is logged with the Feature Store, the transformations are packaged with it, ensuring that identical code is applied to raw data at serving time, effectively eliminating training-serving skew in production pipelines.

Why this answer

The Databricks Feature Store acts as a centralized repository for features. By using the same feature definitions and transformation logic for both training and serving, it eliminates the discrepancy known as training-serving skew. This ensures that the features fed to the model at inference time are calculated exactly as they were during training, maintaining consistent performance and model reliability.

Exam trap

Candidates often confuse model registries with feature stores, failing to identify which specific component is responsible for eliminating training-serving skew.

206
MCQmedium

Which method is the most appropriate for logging custom pre-processing logic alongside a model so that it is automatically applied during inference in Databricks?

A.Include pre-processing in the notebook and rely on the serving user to call it.
B.Use the mlflow.pyfunc flavor to wrap the model and transformation logic.
C.Store the transformation logic in a separate repository and import it at inference time.
D.Write the pre-processing code as a SQL view in the Databricks SQL Warehouse.
AnswerB

The pyfunc model flavor allows developers to define a custom inference contract. By bundling the transformation code (e.g., standard scalers or feature encoding) into the model's predict method, you ensure that any input received by the serving endpoint undergoes the exact same processing that the model expects.

Why this answer

The mlflow.pyfunc flavor is designed specifically for this purpose. It allows developers to define a custom class that inherits from PythonModel, where both the pre-processing logic and the model prediction logic reside in the same artifact. This ensures that the exact same transformation steps applied during training are applied during inference, preventing training-serving skew, which is a common source of production model degradation.

Exam trap

Candidates often suggest creating a separate preprocessing script or pipeline. They overlook that the pyfunc flavor is the standard way to package transformation logic directly with the model artifact.

207
MCQmedium

In an automated MLOps workflow, what is the best practice for handling model training failures?

A.Retrying the training job indefinitely.
B.Ignoring the error and deploying the previous model.
C.Logging the error and triggering an automated notification.
D.Deleting the failed experiment run.
AnswerC

This is the best practice. By capturing the error details and alerting the team, you ensure visibility and quick resolution. Automated logging preserves the context of the failure, which is crucial for debugging, while notifications ensure that the right people are aware of the production pipeline status.

Why this answer

Failing fast and sending alerts is the standard practice. Automated pipelines should include error handling that logs the failure cause to MLflow and triggers a notification via email or Slack. This ensures that data scientists can address the issue immediately without waiting for a manual audit of the pipeline, maintaining the overall health and reliability of the automated machine learning lifecycle.

Exam trap

Candidates often choose manual monitoring or retrying the training job immediately without logging, failing to recognize that automated pipelines require proactive error reporting and systemic logging for root cause analysis.

208
MCQmedium

Which of the following is an advantage of using Databricks AutoML compared to building a custom Scikit-Learn training loop?

A.AutoML models are guaranteed to have higher accuracy than custom models.
B.It generates the training notebook for the trial, allowing for further refinement and transparency.
C.It prevents the need for any feature engineering by automatically creating all required variables.
D.It eliminates the need for monitoring the model after deployment.
AnswerB

Transparency is crucial for MLOps. AutoML produces a fully documented, editable notebook that demonstrates how the model was trained. This allows developers to see the exact preprocessing steps and hyperparameters, giving them full control to refine, optimize, or audit the model logic before deploying it to a production environment.

Why this answer

Databricks AutoML significantly accelerates the development lifecycle by automating tedious tasks like data cleaning, feature engineering, and model selection. It provides a baseline of high-performance models while simultaneously producing the source code for the best-performing model. This allows teams to iterate rapidly, then customize the generated code, combining the speed of automation with the flexibility of manual tuning for complex production requirements.

Exam trap

Candidates often assume AutoML is a complete black box, missing the critical advantage that it outputs the actual training notebook for total code transparency and manual tuning.

209
MCQhard

A machine learning engineer has deployed a model to a Databricks Model Serving endpoint. The model requires a custom Python package that is not available in the default environment. The engineer has already logged the model with MLflow and included the package in the conda environment. However, upon deployment, the endpoint fails to start. What is the most likely cause?

A.The model's MLflow flavor is not supported by Model Serving.
B.The custom package is not available in a public PyPI repository, and no additional index URL was specified.
C.The custom package is not installed on the cluster used for serving.
D.The endpoint's workload size is too small to install the package.
AnswerB

Databricks Model Serving installs dependencies from the conda environment specified in the MLflow model. If the custom package is hosted in a private repository or requires a specific index URL, that must be included in the conda environment's pip section. Without it, the installation fails, causing the endpoint to fail to start. This is a common pitfall when using private packages.

Why this answer

When deploying a model with custom dependencies, Databricks Model Serving uses the conda environment logged with the MLflow model. If the package is not on PyPI or requires a private index, you must specify the index URL in the conda environment. Without it, the installation fails, and the endpoint cannot start.

Ensuring the environment specification includes all necessary sources is critical for successful deployment.

Exam trap

The trap here is assuming that Model Serving automatically has access to all Python packages, but it only installs what is specified in the model's conda environment, including any required index URLs.

210
MCQmedium

You are retiring a real-time model endpoint on Databricks Model Serving. The endpoint has been serving production traffic for six months and you want to archive its request logs for compliance before deleting the endpoint. Which action should you take first?

A.Download the endpoint's model artifact from the MLflow Model Registry and store it as the audit record.
B.Use the Databricks REST API to list the endpoint's configuration and save the JSON response as the archive.
C.Enable inference table logging on the endpoint and wait for the next inference requests to be persisted.
D.Query the existing inference table that was configured for the endpoint and export its contents before deleting the endpoint.
AnswerD

When an endpoint is configured with an inference table, every request and response is written to a Delta table in Unity Catalog. That table persists independently of the endpoint, so querying and exporting it preserves the historical payloads needed for compliance before the endpoint itself is removed.

Why this answer

Model Serving writes inference payloads to a Delta inference table when the feature is enabled, and that table survives endpoint deletion because it is a separate Unity Catalog object. Exporting the table before removing the endpoint preserves the historical request and response records required for compliance, whereas configuration snapshots or model artifacts do not contain traffic data.

Exam trap

The trap here is assuming that enabling inference table logging at deletion time will backfill the historical requests that the endpoint already served.

211
MCQhard

An ML engineer is using MLflow to track a deep learning experiment with PyTorch on Databricks. They want to capture the model's architecture, optimizer state, and training metrics, and later reproduce the exact training run. They call `mlflow.pytorch.autolog()` before training. After several epochs, they notice that metrics are logged but the model signature is missing, and the logged model cannot be loaded for inference without specifying the input example. What should they do to ensure the model is properly logged with a signature?

A.Explicitly log the model using `mlflow.pytorch.log_model()` with the `signature` argument, computed from a sample input using `mlflow.models.infer_signature()`.
B.Use `mlflow.pytorch.save_model()` instead of `log_model()` to automatically include a signature.
C.Provide an input example to `mlflow.pytorch.autolog()` via the `log_every_n_step` parameter.
D.Set the environment variable `MLFLOW_LOG_MODEL_SIGNATURE` to `true` before training.
AnswerA

Autologging for PyTorch may not infer a signature unless an input example is provided. To guarantee a signature, you must explicitly log the model with mlflow.pytorch.log_model and pass a signature created via mlflow.models.infer_signature using a representative input sample. This ensures the model can be loaded for inference without manual input specification and enables validation.

Why this answer

Autologging for PyTorch does not automatically infer a model signature unless an input example is provided. To ensure a signature, explicitly log the model with mlflow.pytorch.log_model and pass a signature generated by mlflow.models.infer_signature using sample input. This enables proper model loading and inference without manual input specification.

Exam trap

The trap here is believing that autologging automatically captures the model signature for PyTorch, when it often requires an explicit input example.

212
MCQmedium

A data scientist is iterating on a model and notices that their training runs are becoming disorganized. What is the standard Databricks mechanism for tracking different 'attempts' at model improvement within a single project?

A.Using a new Git branch for every single hyperparameter change.
B.Grouping runs within an experiment and using MLflow tags.
C.Writing the results to a shared CSV file in the root directory.
D.Storing all run parameters in the notebook's global variable state.
AnswerB

Experiments are the correct organizational container for runs. Tags allow for custom metadata, such as 'production-candidate' or 'v1-feature-set', which makes it easy to filter and search through hundreds of runs to identify the best-performing model, providing a clean and efficient workspace for ongoing research.

Why this answer

MLflow tracking allows for multiple 'runs' under a single experiment. Each run captures parameters, metrics, and artifacts. By systematically naming or tagging these runs, the data scientist can easily compare performance metrics (like RMSE or F1-score) in the MLflow UI.

This structured approach is fundamental to scientific experimentation and prevents the loss of valuable insights when iterating on model architectures or hyperparameter sets.

Exam trap

Candidates often look for complex database solutions or manual folder structures instead of utilizing the built-in experiment grouping and tagging features provided by MLflow for run organization.

213
MCQhard

A machine learning engineer is using Databricks AutoML to train a classification model. They notice that the best model from AutoML has a high F1 score on the validation set but performs poorly on a holdout test set. They suspect that the data has a temporal component and that the default train/validation split is causing data leakage. What should they do to address this?

A.Enable cross-validation with 5 folds in AutoML by setting `n_folds` to 5, which will ensure that each fold respects temporal order.
B.Use `mlflow.log_param` to record the timestamp column and then manually retrain the best model with a custom time-based split outside of AutoML.
C.Re-run AutoML with the `time_col` parameter set to the timestamp column, so that AutoML uses a chronological split for training and validation.
D.Increase the size of the validation set by setting `train_validation_split` to 0.5, which will reduce overfitting and improve generalization.
AnswerC

Databricks AutoML supports a `time_col` parameter that specifies a time column for temporal data. When set, AutoML performs a chronological split, ensuring that training data precedes validation data. This prevents leakage from future data into the training set. It is the correct approach for time-series or temporally ordered data to get realistic validation performance.

Why this answer

Databricks AutoML provides the `time_col` parameter to handle temporal data. When set, AutoML performs a chronological split, ensuring that training data precedes validation data, which prevents leakage from future data. This is the correct way to address the poor holdout performance caused by a random split.

Other options do not fix the temporal leakage issue.

Exam trap

The trap here is assuming that increasing validation size or enabling cross-validation automatically respects temporal order, when only specifying the time column triggers a chronological split.

214
MCQhard

A machine learning engineer is using MLflow to log a custom PyTorch model. They define a custom pyfunc class that inherits from mlflow.pyfunc.PythonModel and implements predict(). After logging the model with mlflow.pyfunc.log_model(), they load it with mlflow.pyfunc.load_model() and call predict() with a pandas DataFrame. The prediction fails with an error about missing context. What is the most likely cause?

A.The model was logged without specifying the conda_env, so the context cannot be reconstructed during loading.
B.The pandas DataFrame passed to predict() must be converted to a NumPy array first, otherwise the context is not passed.
C.The custom pyfunc class must inherit from mlflow.pyfunc.PythonModel and also implement load_context(), which is missing.
D.The predict() method signature must include a context parameter as the first argument, but the implementation omitted it.
AnswerD

In MLflow's PythonModel, the predict() method must accept two arguments: context and model_input. The context provides information about the model and environment. If the implementation defines predict(self, model_input) without context, loading and calling predict will raise an error because MLflow passes the context. This is a common mistake when writing custom pyfunc models.

Why this answer

MLflow's PythonModel.predict() method is defined as predict(self, context, model_input). The context argument is always passed by MLflow when calling predict. If the custom class defines predict(self, model_input), Python will raise a TypeError about missing arguments when MLflow attempts to call it with both context and model_input.

The fix is to include context in the method signature, even if it is not used.

Exam trap

The trap here is assuming that the context parameter is optional or that it relates to environment loading, when it is a required argument in the predict() method signature.

215
MCQhard

A data scientist is using MLflow to log a custom PyTorch model on Databricks. They want to ensure that the model can be loaded and used for inference without requiring the original training code. Which MLflow feature should they use to package the model with its dependencies?

A.`mlflow.pytorch.save_model` to save the model in PyTorch's native format, then log the directory.
B.`mlflow.log_artifact` to save the model's state dict and a requirements.txt file.
C.`mlflow.pytorch.log_model` with the `pip_requirements` parameter to specify dependencies.
D.`mlflow.pyfunc.log_model` with a custom Python function that loads the PyTorch model.
AnswerC

`mlflow.pytorch.log_model` packages the PyTorch model along with a conda environment and pip requirements. By specifying `pip_requirements`, you ensure that all necessary libraries are installed when the model is loaded in a different environment. This makes the model self-contained and reproducible without the original training code.

Why this answer

Using `mlflow.pytorch.log_model` with `pip_requirements` packages the model with its dependencies, creating a self-contained artifact that can be loaded and served without the original training code. Other methods either require manual reconstruction or do not capture dependencies automatically, making them less reliable for deployment.

Exam trap

The trap here is assuming that saving the state dict or using a custom pyfunc wrapper is sufficient, when in fact MLflow's native flavor with explicit dependencies is needed for true portability.

216
MCQhard

Refer to the exhibit. The deployment pipeline is failing to load the model artifact in the target production environment. What is the most likely cause?

A.The model version is not set to 'PRODUCTION' in the metadata.
B.The path '/dbfs/...' is not a reliable way to reference artifacts across different clusters.
C.The conda_env file is missing from the artifact storage.
D.The run_id 'd7a8e9f0' has been archived and is no longer accessible.
AnswerB

Hard-coding DBFS paths is an anti-pattern. MLflow models should be referenced using the 'runs:/' or 'models:/' URI formats. These URIs are managed by the MLflow client and resolve correctly to the underlying artifact location, avoiding the path-resolution issues that occur when using literal file system paths.

Why this answer

The exhibit shows a hard-coded path beginning with '/dbfs/'. In Databricks, accessing DBFS via local file system paths can be inconsistent across different clusters or environments. The correct approach is to use the MLflow URI (e.g., 'runs:/...') to load models, as it abstracts the underlying storage location and ensures the artifact is resolved correctly regardless of the execution environment or storage configuration.

Exam trap

Many test-takers assume local or absolute DBFS paths are safe to use in deployment code, failing to recognize that MLflow URIs are required for environment portability.

217
MCQeasy

A team wants a scheduled Databricks job to automatically retrain a demand-forecasting model whenever the upstream feature table receives new data, and to register the resulting model version only if validation metrics improve. Which Databricks capability should they use to orchestrate this?

A.A cron expression on the training notebook that runs every minute to poll for new feature data.
B.A Delta Live Tables pipeline that materializes the feature table and automatically promotes models with the highest accuracy.
C.An MLflow webhook that fires when a new experiment run is created and calls the registration API.
D.A Databricks job with a table-update trigger that runs the training notebook and a conditional registration task.
AnswerD

Databricks jobs support triggers based on upstream table updates, which is exactly the event described. Chaining a training task with a task that conditionally registers the model based on metrics implements the gated promotion. This uses native orchestration, so no external scheduler or custom polling code is required.

Why this answer

Databricks jobs support table-update triggers, letting a run start when the feature table changes. A multi-task job can then train the model and use a conditional task to register the version only when validation metrics beat the current production baseline, all within native orchestration.

Exam trap

The trap here is reaching for MLflow webhooks or Delta Live Tables for event-driven retraining, when the table-update trigger on a Databricks job is the native mechanism.

218
MCQmedium

A data scientist has registered a model in Unity Catalog under the name `prod.ml.forecast_model`. They now need to define a service-level objective (SLO) that automatically monitors the model's prediction quality in production, not just endpoint uptime. Which Databricks feature should they configure to detect model performance degradation?

A.MLflow Model Registry webhooks
B.Cluster autoscaling policies on the serving cluster
C.Model Serving endpoint health checks
D.Lakehouse Monitoring on the inference table
AnswerD

Lakehouse Monitoring is Databricks' native solution for tracking data and model quality drift. By creating a monitor on the Delta inference table that stores model inputs and predictions, you get automatic profile and drift metrics, including model quality metrics when ground truth is joined. This directly satisfies the requirement to monitor prediction quality, not just infrastructure health.

Why this answer

Lakehouse Monitoring on the inference table is the only option that provides statistical monitoring of model inputs and outputs. It computes drift and profile metrics and, when ground truth labels are available, model quality metrics such as accuracy or RMSE. This makes it the correct mechanism for an SLO tied to prediction performance rather than simple endpoint availability.

Exam trap

The trap here is assuming that serving endpoint health checks or autoscaling provide model quality monitoring, when they only cover infrastructure availability and capacity.

219
MCQeasy

Which feature in Databricks allows you to automatically track training code, parameters, metrics, and models during the development phase?

A.Databricks Delta Lake
B.MLflow Tracking
C.Databricks Feature Store
D.Databricks Workflows
AnswerB

MLflow Tracking provides the API and UI to log training parameters, metrics, and code versions. It is specifically designed to manage the experimental phase of machine learning, allowing data scientists to organize runs, compare performance metrics, and keep track of model artifacts for later deployment in the Model Registry.

Why this answer

MLflow Tracking is the core component of Databricks for experiment management. It provides a structured way to log and visualize the results of machine learning runs. Understanding this component is fundamental to MLOps because it allows teams to compare models, track performance over time, and ensure that every experiment is reproducible, which is essential for professional model development and auditing.

Exam trap

Candidates frequently confuse MLflow Tracking with the MLflow Model Registry. Tracking is for experiment metadata and metrics during development, while the Registry is for lifecycle management and versioning.

220
MCQeasy

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. The model version is 3 and the model name is 'fraud_model'. Which identifier should be used to reference this model version when creating the endpoint?

A.The model's unique ID from the MLflow model registry, such as 'd2f1a3b4c5e6'.
B.The full three-level name and version in the format 'catalog.schema.fraud_model/3'.
C.The full three-level name 'catalog.schema.fraud_model' and the version number '3' specified separately.
D.The model name and version in the format 'fraud_model/3'.
AnswerC

In Unity Catalog, models are referenced by their three-level namespace: catalog.schema.model. The version is a separate integer. When creating a serving endpoint, you provide the model name and the version as distinct fields. This ensures the correct model version is loaded. This option correctly separates the name and version, which is the required format for Databricks Model Serving with Unity Catalog models.

Why this answer

Databricks Model Serving requires the model name and version to be specified separately. For Unity Catalog models, the name is the three-level namespace (catalog.schema.model), and the version is an integer. This allows the serving infrastructure to resolve the exact model version.

Other formats like slash-separated or using the MLflow model version ID are not supported for endpoint creation.

Exam trap

The trap here is assuming that the model version is part of the model name or that a unique ID can be used, when in fact the name and version must be provided as separate fields in the endpoint configuration.

221
MCQmedium

Which Databricks feature is specifically designed to manage the lifecycle of a machine learning model, including versioning, stage transitions, and deployment tracking?

A.Unity Catalog.
B.MLflow Model Registry.
C.Databricks Delta Lake.
D.Databricks Repos.
AnswerB

The Model Registry is built exactly for the needs of model lifecycle management. It provides a centralized hub to track model versions, handle approvals, and manage deployments, which is essential for ensuring that ML models in production are stable, reproducible, and compliant with organizational standards for model deployment and auditing.

Why this answer

The MLflow Model Registry is the definitive tool in Databricks for managing the model lifecycle. It allows teams to register models, track versions, and manage stage transitions (e.g., Staging to Production). By providing a centralized, audit-trailed repository, it ensures that only validated models are deployed into production, fulfilling the core requirements of MLOps for governance, reliability, and automated deployment pipelines.

Exam trap

Candidates confuse MLflow Tracking with the MLflow Model Registry, failing to realize that governance, versioning, and stage transitions are handled by the Registry.

222
MCQmedium

A data science team trains a scikit-learn model on a Databricks cluster and needs the same feature-engineering logic to run identically in a nightly batch scoring job and in a real-time Model Serving endpoint. They want a single artifact that encapsulates preprocessing and the estimator. Which approach should they use?

A.Log the model with MLflow using the sklearn flavor and a custom pyfunc wrapper that includes the preprocessing steps.
B.Register the raw estimator in Unity Catalog and reimplement preprocessing separately in the batch job and endpoint code.
C.Package the preprocessing into a custom container image and reference the image from the model version metadata.
D.Log the model with the sklearn flavor and pass the preprocessing function name as a signature parameter.
AnswerA

The pyfunc flavor wraps arbitrary Python logic, so preprocessing and the estimator travel as one artifact. Batch jobs load it with mlflow.pyfunc.load_model and the serving endpoint loads the same artifact, guaranteeing identical transformations. This is the standard Databricks pattern for eliminating training-serving skew when feature logic must be shared.

Why this answer

Wrapping preprocessing and the estimator in a single pyfunc model lets MLflow serialize both as one artifact. Batch scoring loads that artifact with the pyfunc loader, and Model Serving uses the same artifact, so feature transformations are guaranteed identical and training-serving skew is avoided without duplicating logic.

Exam trap

The trap here is believing a model signature or container image reference can carry executable preprocessing, when only a pyfunc wrapper actually bundles that logic into the artifact.

223
MCQmedium

Which TWO statements regarding the use of Unity Catalog in Databricks for MLOps are correct?

A.Unity Catalog enforces access control at the table level only, not on MLflow models.
B.It enables centralized lineage tracking between data tables and MLflow models.
C.Unity Catalog is only compatible with Python and does not support SQL-based model management.
D.Models registered in Unity Catalog can be shared across multiple workspaces.
E.Users must manually export model artifacts to an external store to share them via Unity Catalog.
AnswerB, D

Unity Catalog provides a unified lineage graph that shows how data flows from source tables into ML models. This visibility is essential for understanding the impact of data changes on model performance and for fulfilling regulatory requirements regarding data provenance and model transparency.

Why this answer

Unity Catalog acts as the centralized governance layer for all data and AI assets. It provides fine-grained access control and lineage tracking, allowing organizations to maintain visibility over which data is used by which models. By centralizing these assets, teams can collaborate safely while ensuring that compliance and auditing requirements are satisfied across the entire organization, regardless of the individual workspace where the work is performed.

Exam trap

Candidates often assume that models registered in Unity Catalog are isolated to a single workspace, missing that cross-workspace sharing and centralized governance are core advantages of using Unity Catalog for enterprise ML.

224
MCQhard

When using MLflow to manage the machine learning lifecycle, what is the primary purpose of the 'conda.yaml' or 'requirements.txt' file automatically generated during log_model?

A.To store the model's hyperparameter search space configurations.
B.To ensure that the inference environment has the necessary dependencies installed.
C.To act as a security manifest that validates the digital signature of the model.
D.To limit the maximum number of concurrent requests the model can handle.
AnswerB

The environment file acts as a manifest for the model. By documenting the exact versions of all libraries used, it allows the deployment target to reconstruct the runtime accurately. This is the key mechanism that enables 'write once, deploy anywhere' functionality within the MLflow lifecycle.

Why this answer

These files define the environment specification required to recreate the model's runtime environment. When a model is moved to a production serving endpoint or a different cluster, Databricks uses these specifications to install the correct library versions. This ensures that the model executes in an environment identical to the one it was trained in, preventing silent failures caused by library version mismatches.

Exam trap

Candidates assume automatically generated environment files are only for documentation, ignoring their critical role in setting up exact dependencies for production inference.

225
MCQhard

A team is building an automated retraining pipeline. They need to ensure that only models exceeding a certain performance threshold are registered. What is the most effective way to implement this logic?

A.Manually inspect all models in the MLflow UI before registering them.
B.Write a Python script that evaluates the model and conditionally calls the Model Registry API.
C.Register every trained model and let the production system filter them.
D.Use MLflow's 'auto-register' feature that registers all models by default.
AnswerB

Using a script to programmatically evaluate the model and trigger registration based on performance metrics creates a robust gate. This ensures consistent, reproducible, and automated quality control, allowing the team to maintain high performance standards without manual oversight, which is necessary for modern, efficient, and scalable machine learning production workflows.

Why this answer

Integrating conditional logic into the retraining script using the MLflow API allows for programmatic model governance. By evaluating the model against validation data and only calling 'register_model' if the performance exceeds the threshold, the team prevents poor-quality models from entering the registry. This automated gatekeeping is vital for maintaining the health of the production pipeline and ensuring that only high-performing models proceed to deployment.

Exam trap

Candidates often assume the Model Registry has built-in automatic threshold triggers, leading them to select incorrect answers that imply a configuration setting rather than programmatic API implementation.

Page 2

Page 3 of 4

Page 4

All pages

Practice Databricks-ML-Pro by domain

Target a specific domain to shore up weak areas.

See all domains with question counts →