Courseiva

Databricks Certified Machine Learning Professional (Databricks-ML-Pro) — Questions 76–150

300 questions total · 4pages · All types, answers revealed

Page 1

Page 2 of 4

Page 3
76
MCQmedium

A machine learning engineer is using MLflow to track experiments and has logged a model with a signature. They now want to register this model in the MLflow Model Registry and promote it to Production. Which MLflow API call should they use to add the model to the registry?

A.mlflow.log_model(model, artifact_path, registered_model_name=name)
B.mlflow.pyfunc.save_model(path, loader_module, data_path, registered_model_name=name)
C.mlflow.register_model(model_uri, name)
D.mlflow.create_registered_model(name)
AnswerC

mlflow.register_model is the correct API to register an existing model artifact from a run into the Model Registry. It takes the model URI (e.g., runs:/<run_id>/model) and the desired model name. This call creates a new model version in the registry, which can then be transitioned to Production.

Why this answer

The mlflow.register_model function is designed to take an existing model URI and add it as a new version to the Model Registry. It is the standard way to register a model after logging. Other options either create empty models or log new models, which do not fit the scenario of promoting an already logged model.

Exam trap

The trap here is confusing the API for creating a registered model container with the API for adding a version from an existing run.

77
MCQhard

A machine learning engineer is setting up a CI/CD pipeline for a model deployed to Databricks Model Serving. The pipeline must automatically update the serving endpoint when a new model version is registered in the MLflow Model Registry and passes a validation job. Which Databricks feature should the engineer use to trigger the update?

A.Databricks Model Serving's built-in auto-update feature, which automatically deploys the latest model version from the registry.
B.MLflow Model Registry webhooks that call a Databricks job to validate and update the endpoint.
C.A Databricks notebook with a while loop that continuously checks the Model Registry for new versions and updates the endpoint.
D.Databricks Jobs with a schedule trigger that polls the Model Registry every hour for new versions.
AnswerB

Registry webhooks can trigger on MODEL_VERSION_CREATED events, invoking a Databricks job that runs validation and then updates the serving endpoint via the REST API. This provides an event-driven, automated pipeline that meets the requirement of updating upon registration and successful validation.

Why this answer

MLflow Model Registry webhooks enable event-driven automation. By configuring a webhook on MODEL_VERSION_CREATED, a Databricks job can be triggered to validate the new version and then update the Model Serving endpoint via the REST API. This integrates validation gating and deployment in a CI/CD pipeline.

Exam trap

The trap here is assuming Model Serving has an auto-update feature or that polling is sufficient, when event-driven webhooks are the intended mechanism for registry-triggered deployments.

78
MCQmedium

A machine learning engineer is training a model using MLflow on Databricks and wants to ensure that the model's input schema is captured and enforced during inference. They are using the `mlflow.pyfunc` flavor. Which action should they take to enable schema enforcement?

A.Log the model with the `signature` parameter, specifying the input and output schema.
B.Save the schema as a separate artifact and load it manually in the inference code.
C.Use `mlflow.set_tag` to record the schema as a JSON string in the run's tags.
D.Enable autologging for the specific framework, which automatically captures and enforces the schema.
AnswerA

Providing a signature when logging the model with `mlflow.pyfunc.log_model` captures the expected input and output schema. This signature is used during inference to validate that the input data matches the expected schema, helping to prevent errors and ensuring consistency. It is the standard way to enable schema enforcement in MLflow.

Why this answer

The correct way to enable schema enforcement is to log the model with a signature. This signature is stored with the model and used by MLflow's serving components to validate incoming data. Tags and artifacts do not provide runtime enforcement, and autologging does not enforce schemas during inference.

Exam trap

The trap here is confusing schema documentation (tags or artifacts) with schema enforcement (signature).

79
MCQmedium

You are training a model on Databricks using MLflow. You need to log a custom model flavor to ensure it can be loaded in an environment without the original training code. Which approach is best practice?

A.Serialize the model object using pickle and upload it to DBFS.
B.Use mlflow.sklearn.log_model with a hardcoded path to the local file system.
C.Create a Python class inheriting from PythonModel and log it via mlflow.pyfunc.log_model.
D.Register the model directly using the Model Registry without logging it to an experiment run.
AnswerC

Inheriting from PythonModel provides a standardized way to define custom predict methods. This ensures that the model can be loaded by any client with the pyfunc flavor, regardless of the underlying library used for training, fulfilling the requirement for environment-agnostic deployment in Databricks.

Why this answer

Using the mlflow.pyfunc.log_model method allows you to encapsulate model artifacts, environment dependencies, and custom inference logic into a single, portable object. This is critical for Databricks model deployment because it decouples the inference environment from the training notebook, ensuring reproducibility across different clusters or serving endpoints while maintaining a standardized interface for prediction.

Exam trap

Candidates often suggest pickling the model object directly or saving it as a raw artifact, which fails to capture the necessary dependency environment, making the model non-portable across different clusters.

80
MCQeasy

When training a model in Databricks, which storage layer should you prioritize for training data to ensure maximum throughput and compatibility with Feature Store?

A.CSV files stored in DBFS.
B.Delta tables.
C.Local temporary storage on the driver node.
D.JSON files stored in an unmanaged external mount.
AnswerB

Delta tables are optimized for high-performance reading and writing in Databricks. They support the versioning and time-travel features required by Feature Store, which ensures that training sets can be recreated precisely, promoting model reproducibility and enabling seamless integration with the broader Databricks ML ecosystem.

Why this answer

Delta Lake is the recommended storage format for training data in Databricks. It provides ACID transactions, schema enforcement, and efficient data versioning, which are essential for reproducible machine learning. Using Delta tables allows Feature Store to perform time-travel queries, ensuring that the data used for training is consistent with the data used for inference, which minimizes the risk of data leakage.

Exam trap

Candidates often recommend standard parquet files or raw object storage, forgetting that Delta tables provide the essential transactional versioning and time-travel required by Feature Store.

81
MCQeasy

A machine learning engineer is setting up a Databricks job to retrain a model every night. The job must use a specific Python library version that is not available in the default Databricks Runtime. The engineer wants to ensure that the retraining job has access to this library without affecting other jobs in the workspace. What is the recommended approach?

A.Modify the global init script to install the library on all clusters in the workspace.
B.Install the library on the job cluster by adding it as a cluster library, and configure the job to use that cluster.
C.Create a custom Docker image with the required library and configure the job cluster to use that image.
D.Install the library using `%pip install` in the first cell of the notebook used by the job.
AnswerB

In Databricks, you can install libraries at the cluster level, making them available to all notebooks and jobs that run on that cluster. For a job, you can define a job cluster with the required library installed. This ensures the library is present for every run without manual installation, and it isolates the library to that cluster, preventing impact on other jobs. This is the recommended and simplest approach for job-specific dependencies.

Why this answer

For a Databricks job that requires a specific library, the best practice is to install the library on the job cluster. Job clusters are dedicated to the job and can be configured with libraries that are installed at cluster startup. This ensures the library is available for every run, isolates dependencies, and avoids impacting other workloads.

It is simpler and more maintainable than custom Docker images or notebook-level installations.

Exam trap

The trap here is thinking that installing a library in a notebook with `%pip install` is sufficient for a scheduled job, but that installation is ephemeral and not isolated, so it fails to provide a persistent, job-specific environment.

82
MCQeasy

Which Databricks feature is specifically designed to prevent data leakage during model training by ensuring feature values are fetched as they existed at a specific point in time?

A.Delta Lake Time Travel.
B.Databricks Feature Store.
C.MLflow Model Registry.
D.Auto Loader.
AnswerB

The Feature Store automatically handles point-in-time joins using specified lookup keys and timestamps. This functionality is essential for preventing data leakage during training, as it ensures that the training dataset only contains information that would have been available at the moment of prediction in a real-world setting.

Why this answer

Point-in-time joins are the core mechanism of the Databricks Feature Store. When training on historical data, it is critical to use features that were available at that time, not their current values. This prevents 'look-ahead bias,' where future information inadvertently leaks into the training set, causing the model to perform artificially well during training but failing in production.

Exam trap

Students frequently confuse standard data caching or cross-validation with point-in-time correctness, missing the specialized purpose of the Databricks Feature Store.

83
MCQhard

A machine learning engineer is building a model on Databricks and wants to use MLflow to track experiments. They need to log a custom metric that is calculated during training but is not automatically captured by `mlflow.autolog()`. They also want to ensure that the metric is associated with the correct run. Which code snippet should they use inside their training script?

A.`mlflow.log_metric("custom_metric", value)` after starting a run with `mlflow.start_run()`.
B.`mlflow.set_tag("custom_metric", value)` after starting a run with `mlflow.start_run()`.
C.`mlflow.log_artifact("custom_metric", value)` after starting a run with `mlflow.start_run()`.
D.`mlflow.log_param("custom_metric", value)` after starting a run with `mlflow.start_run()`.
AnswerA

`mlflow.log_metric` is the correct API to log a custom metric manually. It must be called within an active MLflow run, typically started with `mlflow.start_run()`. This ensures the metric is associated with that specific run. This approach is standard for logging metrics not captured by autologging, such as custom evaluation scores or business-specific KPIs.

Why this answer

To log a custom metric in MLflow, the correct API is `mlflow.log_metric`, which records a scalar value associated with the current run. It must be called within an active run context. This allows the metric to be tracked, visualized, and compared across runs.

Other APIs like `log_param`, `set_tag`, or `log_artifact` serve different purposes and would not properly record the metric.

Exam trap

The trap here is confusing the different MLflow logging APIs and using one meant for parameters, tags, or artifacts instead of metrics.

84
MCQhard

You are debugging a model serving issue where the model is failing in production. Which THREE of the following actions should you prioritize to identify the root cause of the failure?

A.Check the Databricks cluster/endpoint logs for errors during the inference request.
B.Review the input data schema and compare it with the training data expectations.
C.Retrain the model from scratch on the full production dataset to clear errors.
D.Examine the model lineage in MLflow to identify the exact code and data used.
E.Delete the model registry and recreate the model versions.
AnswerA, B, D

Logs are the primary source of truth for runtime failures. They will contain stack traces if the model code crashed, memory exhaustion errors, or connectivity issues with external services. Reviewing logs is the fastest way to determine if the issue is environmental or logic-based.

Why this answer

Effective troubleshooting in MLOps relies on systematic isolation of the failure point: infrastructure, data, or model code. Checking system logs, verifying input data schemas, and reviewing model version history are standard diagnostic steps. These actions help determine if the failure is due to a sudden change in incoming data distributions, an infrastructure error, or a regression within the specific model version currently deployed to production.

85
Multi-Selecthard

You are implementing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically run unit tests, train the model, and deploy it to a staging endpoint. Which TWO practices should you follow to ensure the pipeline is reproducible and reliable? (Choose two.)

Select 2 answers
A.Use the latest version of all libraries to ensure compatibility and security patches.
B.Pin all library dependencies to specific versions in a requirements file or conda environment specification.
C.Configure the pipeline to run on a cluster with autoscaling enabled to handle variable workloads.
D.Manually promote the model to staging after reviewing the test results to ensure quality.
E.Store all code, including notebooks and Python modules, in a Git repository and use Databricks Repos to check out the code in the pipeline.
AnswersB, E

Pinning dependencies ensures that the environment is consistent across runs, preventing issues caused by library updates. This is critical for reproducibility in ML pipelines, where even minor version changes can affect model training and inference. Using a requirements file or conda environment specification allows the pipeline to recreate the exact environment.

Why this answer

For a reproducible and reliable CI/CD pipeline, code should be version-controlled and checked out consistently, and dependencies should be pinned to specific versions. These practices ensure that the pipeline runs the same code in the same environment every time. Autoscaling and manual promotion do not address reproducibility, and using latest libraries can introduce instability.

Exam trap

The trap here is assuming that using the latest libraries or manual promotion improves reliability, when in fact they introduce variability and reduce automation.

86
MCQmedium

A fraud detection model is served via a Databricks Model Serving endpoint. The team wants to capture the incoming request payloads and the model's predictions to a Delta table for monitoring and future retraining. Which approach is most appropriate?

A.Enable inference tables on the model serving endpoint to automatically log requests and responses to a Delta table.
B.Modify the model's predict method to write each request and response to a Delta table using the Spark connector.
C.Configure the endpoint to send logs to a cloud storage bucket and schedule a job to load them into Delta.
D.Use MLflow autologging on the serving endpoint to capture inference requests as MLflow runs.
AnswerA

Inference tables are a Databricks Model Serving feature that captures request and response payloads to a Delta table automatically. This provides the exact logging needed for monitoring and retraining without modifying the model code or building custom logging infrastructure.

Why this answer

Databricks Model Serving provides inference tables as a built-in mechanism to log request and response payloads to a Delta table. This satisfies monitoring and retraining data capture without custom code. The other options either add latency, require unsupported dependencies in the serving environment, or confuse training-time autologging with serving-time logging.

Exam trap

The trap here is assuming you must instrument the model code or build a custom pipeline, when Model Serving already offers inference tables for request and response logging.

87
MCQmedium

A machine learning engineer is training a scikit-learn model on Databricks and wants to automatically log hyperparameters, metrics, and the trained artifact without writing extensive boilerplate logging code. Which approach should the engineer use?

A.Manually invoke mlflow.start_run() and explicitly call mlflow.log_metric() for every single training epoch iteration.
B.Execute mlflow.sklearn.autolog() prior to initiating the model training fit process to automatically capture all runs.
C.Store all model metrics directly inside a temporary Delta table and sync them to MLflow during the final evaluation phase.
D.Configure the cluster environment variables to automatically intercept stdout and parse training metrics using regex.
AnswerB

Calling mlflow.sklearn.autolog() before fit() hooks into scikit-learn's training APIs, capturing hyperparameters, metrics, and the fitted model artifact automatically, with no manual logging calls. This directly satisfies the stem's requirement to avoid extensive boilerplate logging code during Databricks training runs.

Why this answer

Autologging automatically captures all relevant parameters, metrics, and models without requiring manual mlflow.log_param calls. This ensures consistency, prevents human error during tracking, and natively integrates with Spark dataframes. It significantly accelerates the model development lifecycle on Databricks by standardizing the capture of experimentation metadata across distributed runs.

Exam trap

Candidates mistakenly call autolog() after the training fit process or confuse it with manual logging functions, missing that it must be initialized beforehand to capture all metrics.

88
MCQmedium

A data scientist is deploying a model to Databricks Model Serving. The model was trained using a scikit-learn pipeline that includes a custom transformer. The custom transformer is defined in a Python module that is not part of the model artifact. What should the data scientist do to ensure the model can be served successfully?

A.Include the custom transformer code in the model's conda environment as a pip dependency.
B.Log the model with the custom transformer code included in the model's artifacts using MLflow's custom logging.
C.Re-train the model without the custom transformer, as Model Serving does not support custom code.
D.Convert the custom transformer to a built-in scikit-learn transformer before logging.
AnswerB

MLflow allows you to log custom code along with the model by including it in the model's artifacts and specifying it in the model's Python function or by using the `code_path` parameter in `mlflow.pyfunc.log_model`. This ensures the custom transformer is available at inference time. Databricks Model Serving will then load the code from the model artifact, making it the correct and straightforward solution.

Why this answer

To deploy a model with custom code, you must include that code in the model's artifacts when logging with MLflow. Using the `code_path` parameter in `mlflow.pyfunc.log_model` allows you to specify additional code files that will be packaged with the model. Databricks Model Serving then loads this code, ensuring the custom transformer is available during inference.

This is the standard method for handling custom logic in served models.

Exam trap

The trap here is thinking that custom code must be installed as a separate package; instead, it can be bundled directly with the model artifact using MLflow's code_path.

89
MCQeasy

A data scientist is using MLflow to track experiments in a Databricks notebook. They want to record the source code version (Git commit hash) automatically with each run. Which MLflow feature should they enable to capture this information?

A.Enable MLflow's Git integration by setting the `MLFLOW_GIT_COMMIT` environment variable before starting the run.
B.Ensure the notebook or script is running from a Git repository, and MLflow will automatically log the Git commit hash as a tag on the run.
C.Set the `MLFLOW_TRACKING_URI` environment variable to point to a Git repository.
D.Use `mlflow.start_run()` with the `tags` parameter to manually log the Git commit hash.
AnswerB

MLflow automatically captures Git metadata, including the commit hash, repository URL, and branch, when the code is executed from within a Git repository. In Databricks, if the notebook is part of a Git folder (Repos), MLflow will log the Git commit hash as a tag on the run. This is an automatic feature and requires no additional configuration beyond being in a Git repo. This is the correct way to capture source code version.

Why this answer

MLflow has built-in Git integration that automatically logs the Git commit hash, repository URL, and branch as tags when the code is run from a Git repository. In Databricks, this works when using Databricks Repos. No manual tagging or environment variables are needed.

The other options either reference incorrect environment variables or require manual intervention, which does not meet the requirement for automatic capture.

Exam trap

The trap here is thinking that an environment variable like `MLFLOW_GIT_COMMIT` must be set, when MLflow automatically captures Git information if the code is in a Git repository.

90
MCQmedium

Your organization needs to automate the deployment of models to real-time serving endpoints. Which service within the Databricks ecosystem handles the hosting and scaling of these endpoints with managed containerization?

A.Databricks Jobs
B.Databricks Model Serving
C.Delta Live Tables
D.Databricks Repos
AnswerB

Model Serving is the dedicated service for hosting models. It manages the server infrastructure, auto-scaling, and deployment of models from the registry. This allows data science teams to deploy models via a single click or API call, ensuring consistent, scalable production performance without manual configuration of servers.

Why this answer

Databricks Model Serving provides a fully managed, scalable infrastructure for deploying models as REST APIs. It handles the underlying container orchestration, allowing developers to focus on the model artifact rather than infrastructure management. This is essential for MLOps, as it simplifies the transition from a registered model in the MLflow registry to a production-ready, highly available endpoint that handles real-time traffic efficiently.

Exam trap

Candidates often confuse Databricks Model Serving with MLflow tracking or model registry features, picking general tracking tools instead of the dedicated endpoint hosting service designed for real-time containerized serving.

91
MCQmedium

Your team is migrating models to Databricks Model Registry. You need to automate the transition of a model version to 'Staging' only after it passes an automated integration test suite in the CI/CD pipeline. Which mechanism should you use to best achieve this?

A.Manually update the model stage using the Databricks UI after the tests complete.
B.Configure a Databricks Job to transition the model based on a scheduled cron trigger.
C.Use the MLflow Python client's transition_model_version_stage method within the CI/CD script.
D.Set the model stage to 'Staging' inside the training notebook using mlflow.register_model.
AnswerC

The 'transition_model_version_stage' method is the standard programmatic way to manage lifecycle transitions in Databricks. By embedding this in the pipeline, the system verifies test results first, then immediately updates the stage, ensuring that only validated models are marked as ready for further staging or production deployment.

Why this answer

Integrating the MLflow client within the CI/CD pipeline allows for programmatic stage transitions using the 'transition_model_version_stage' API. This approach ensures that manual intervention is minimized and models only move forward after passing defined quality gates. Automating this lifecycle phase is a cornerstone of robust MLOps, as it prevents untested models from reaching production environments and ensures consistency across deployment stages.

92
MCQmedium

An ML engineer is updating a model serving endpoint to use a new model version. They want to gradually shift traffic from the old version to the new version to monitor performance before full rollout. Which feature of Databricks Model Serving should they use?

A.Configure traffic splitting on the existing endpoint by specifying a percentage of traffic to route to each served model version.
B.Create a new endpoint for the new model version and use a load balancer to distribute traffic.
C.Use the 'Champion' and 'Challenger' aliases in Unity Catalog to automatically route traffic based on model performance.
D.Enable A/B testing by deploying the new model to a separate endpoint and using Databricks SQL to compare metrics.
AnswerA

Databricks Model Serving supports traffic splitting, allowing you to route a percentage of requests to different model versions within the same endpoint. By specifying the traffic percentage for each served entity, you can gradually shift traffic from the old to the new version. This enables safe rollout and performance monitoring without disrupting the endpoint.

Why this answer

Databricks Model Serving allows configuring traffic splitting on a single endpoint by assigning a percentage of traffic to each served model version. This enables gradual rollout, allowing the team to monitor the new version's performance while still serving the old version. It is the native and recommended approach for safe model updates.

Exam trap

The trap here is thinking that aliases or separate endpoints automatically handle gradual traffic shifting, when in fact traffic splitting must be explicitly configured on the endpoint.

93
Multi-Selecthard

A team is deploying a model that requires custom Python libraries not available in the default Databricks Runtime. Which TWO methods can be used to ensure these dependencies are available in the Model Serving environment? (Select TWO)

Select 2 answers
A.Include a 'requirements.txt' file or a 'conda.yaml' when calling mlflow.log_model().
B.Manually install the libraries on the driver node of the cluster used for training.
C.Use the pip_requirements parameter in the mlflow.sklearn.log_model (or similar) function.
D.Add the libraries to the Spark configuration of the Model Serving endpoint.
E.Upload the wheel files to a DBFS location and reference them in the endpoint UI.
AnswersA, C

Providing a requirements file or Conda environment during logging tells MLflow exactly which packages and versions are needed. When the model is deployed to an endpoint, Databricks uses this information to build a container image with all the necessary dependencies pre-installed and ready for execution.

Why this answer

Model Serving environments are reconstructed based on the metadata captured when the model was logged. To include custom libraries, they must be explicitly defined during the logging process. This ensures that the production environment exactly matches the development environment, preventing 'missing module' errors during real-time inference requests.

Exam trap

Test-takers often think installing libraries directly in the serving endpoint settings or cluster configuration is sufficient, forgetting that dependencies must be captured during model logging.

94
Multi-Selectmedium

An ML engineer is transitioning a model from the Workspace Model Registry to the Unity Catalog (UC) Model Registry. Which TWO statements describe benefits or requirements of using Unity Catalog for model management? (Select TWO)

Select 2 answers
A.Unity Catalog supports model sharing across multiple Databricks workspaces.
B.Models must be stored in a legacy DBFS location to be registered in Unity Catalog.
C.Unity Catalog models use a three-level namespace: catalog, schema, and model name.
D.The legacy MLflow Stages (Staging, Production) are the primary way to manage UC models.
E.Only models logged with the SparkML flavor can be registered in the Unity Catalog.
AnswersA, C

By utilizing a centralized Metastore, Unity Catalog allows models to be registered once and accessed by authorized users across any workspace linked to that Metastore. This facilitates better collaboration and standardization across large organizations that use separate environments for development, testing, and production stages.

Why this answer

Unity Catalog centralizes governance for all data and AI assets, providing a unified interface for access control and lineage. Transitioning to UC enables fine-grained permissions and allows models to be shared across different workspaces within the same Metastore. This is a significant improvement over the legacy workspace-specific registry which lacked cross-workspace visibility and unified auditing.

Exam trap

Candidates often confuse Unity Catalog with legacy workspace registry features, failing to recognize that UC's main value is cross-workspace sharing and the three-level namespace structure.

95
MCQmedium

When deploying a model to a production environment, why is it recommended to use a 'Model Signature'?

A.To increase the inference speed of the model.
B.To define the input and output schema for automatic validation.
C.To automatically retrain the model if the input data changes.
D.To encrypt the model artifacts in the registry.
AnswerB

The signature provides a contract for the model, detailing the expected data types and shapes. The serving infrastructure uses this contract to validate every API request, catching mismatches before they reach the model code. This prevents crashes and provides clear error messages when invalid data is submitted.

Why this answer

A model signature defines the schema of the inputs and outputs. By enforcing this, the serving endpoint can automatically validate incoming requests and reject those that do not match the expected format. This prevents runtime errors and unexpected behavior, serving as a critical safety feature in automated production pipelines where data quality might fluctuate, ensuring the model always receives the correct data structure.

Exam trap

Candidates often confuse model signatures with performance metrics or hyperparameters, failing to realize that signatures specifically define the expected input and output data schemas.

96
MCQhard

A machine learning engineer is using Hyperopt with SparkTrials on a Databricks cluster to tune a gradient boosting model. They notice that the tuning job is running slowly because each trial trains on the full dataset, and they want to speed up the search without sacrificing final model quality. Which approach is most appropriate?

A.Use a smaller subset of the training data for each trial and then retrain the best model on the full dataset.
B.Increase the parallelism parameter in SparkTrials to run more trials concurrently.
C.Reduce the number of trials and increase the max_evals parameter to compensate.
D.Switch from SparkTrials to Trials to avoid distributed overhead.
AnswerA

Training on a subset for hyperparameter search reduces the time per trial, allowing more configurations to be explored quickly. After identifying the best hyperparameters, retraining on the full dataset recovers full model quality. This is a common and effective strategy when the dataset is large and individual trials are slow, as it balances search efficiency with final performance.

Why this answer

Using a data subset for hyperparameter tuning accelerates the search by reducing training time per trial, enabling more configurations to be evaluated. Once the best hyperparameters are found, retraining on the full dataset ensures the final model benefits from all available data. This balances efficiency and quality, especially when full-dataset training is expensive.

Exam trap

The trap here is focusing on parallelization or trial count when the bottleneck is per-trial training time due to full dataset usage.

97
MCQhard

A machine learning engineer is developing a custom MLflow Python model that requires a pre-processing step using a scikit-learn pipeline. They want to log the model such that it can be served with the pipeline included. Which approach should they take?

A.Create a Python class that inherits from `mlflow.pyfunc.PythonModel`, include the pipeline in the `predict` method, and log with `mlflow.pyfunc.log_model`.
B.Log the pipeline as a separate model and use MLflow's multi-model serving to chain them.
C.Log the scikit-learn pipeline directly with `mlflow.sklearn.log_model`.
D.Use `mlflow.sklearn.log_model` and pass the pipeline as an artifact, then load it manually in a custom serving script.
AnswerA

Subclassing `mlflow.pyfunc.PythonModel` allows encapsulating custom pre-processing and the scikit-learn pipeline. The `predict` method can call the pipeline and any additional logic. Logging with `mlflow.pyfunc.log_model` saves the model with its dependencies, enabling serving with the full pipeline. This is the standard approach for custom models.

Why this answer

For custom models that include pre-processing and a scikit-learn pipeline, the recommended approach is to create a custom `mlflow.pyfunc.PythonModel` subclass. The `predict` method can invoke the pipeline and any additional logic. Logging with `mlflow.pyfunc.log_model` packages the model and its dependencies for serving.

This ensures the entire pipeline is included.

Exam trap

The trap here is assuming that logging the scikit-learn pipeline alone is sufficient, when custom logic may require a pyfunc wrapper.

98
MCQeasy

A machine learning engineer needs to automate retraining of a model whenever new data lands in a Delta table. The retraining must run on a schedule, use a specific cluster configuration, and send an email alert on failure. Which Databricks feature should be used to orchestrate this workflow?

A.Databricks Jobs with a notebook task, a schedule trigger, and an email notification on failure.
B.An MLflow Project run triggered by a file arrival event in DBFS using a filesystem watcher.
C.A Delta Live Tables pipeline with a continuously running mode that trains the model in a streaming table.
D.A Databricks SQL dashboard with a scheduled refresh that calls a stored procedure to retrain the model.
AnswerA

Databricks Jobs provide scheduling, cluster specification, and notification configuration in one place. A notebook task can read the Delta table, retrain, and register the model. Schedule triggers and failure notifications are built-in, directly meeting all stated requirements without external orchestration tools.

Why this answer

Databricks Jobs is the native orchestration service for scheduled notebooks and scripts. It supports specifying cluster configuration, scheduling, and failure notifications. The other options either lack scheduling, cannot run training code, or are designed for data pipelines rather than ML workflows, so they do not satisfy the combined requirements of schedule, cluster control, and alerting.

Exam trap

The trap here is reaching for specialized ML features like MLflow Projects or Delta Live Tables when the requirements are basic scheduling, cluster control, and alerting that Databricks Jobs already provides.

99
Multi-Selectmedium

You are responsible for monitoring a critical model deployed to Databricks Model Serving. You need to detect data drift and model performance degradation. Which TWO of the following actions should you take? (Choose two.)

Select 2 answers
A.Deploy the model to a different workspace to isolate production traffic.
B.Set up a Databricks SQL dashboard that queries the inference table and compares feature distributions to a baseline.
C.Enable inference tables on the model endpoint to log request and response data.
D.Use MLflow to log the model's training metrics and compare them to production metrics manually.
E.Configure the endpoint to use a smaller instance type to reduce cost.
AnswersB, C

Databricks SQL dashboards can query inference tables to compute statistics and compare them against baseline distributions. This allows you to visualize drift and set up alerts. By leveraging SQL, you can create scheduled queries that detect significant deviations and trigger notifications, providing an effective monitoring solution without additional infrastructure.

Why this answer

Inference tables capture the necessary request and response data for monitoring, while Databricks SQL dashboards enable analysis and alerting on that data. Together, they provide a robust solution for detecting data drift and model performance degradation. Other actions, such as changing instance types or isolating workspaces, do not address the monitoring requirements.

Exam trap

The trap here is focusing on infrastructure changes or manual comparisons instead of leveraging automated data capture and analysis tools designed for monitoring.

100
MCQmedium

A data scientist is training a model on a large Delta table. They want to ensure that the training data remains consistent even if the underlying table is updated during the training process. What is the most robust way to achieve this?

A.Create a temporary local copy of the table in the driver's memory.
B.Use Delta Lake Time Travel to query the table at a specific version.
C.Lock the table using a SQL 'LOCK TABLE' statement during the job.
D.Manually filter data by a 'processed_timestamp' column in every query.
AnswerB

Time travel allows the model training job to reference an immutable snapshot of the Delta table. This ensures that the dataset used for training remains exactly the same, even if new data is appended or existing rows are updated, providing the necessary stability for reliable, reproducible machine learning experiments.

Why this answer

Delta Lake's time travel feature allows users to query a specific version or timestamp of a table. By using `VERSION AS OF` or `TIMESTAMP AS OF` syntax, the training job can anchor itself to an immutable snapshot of the data. This is critical for model reproducibility, as it ensures that the training set does not change mid-experiment, preventing non-deterministic results and enabling accurate auditing.

Exam trap

Candidates often suggest copying the table to a new location. They fail to utilize Delta Lake's native Time Travel feature, which is the standard, efficient method for snapshotting data.

101
MCQmedium

You are responsible for deploying a machine learning model to a Databricks Model Serving endpoint. The model must be updated frequently with new versions. You want to ensure that the endpoint remains available during updates and that traffic can be shifted gradually to the new version. Which deployment strategy should you use?

A.Rolling update by stopping the endpoint, updating the model, and restarting the endpoint.
B.Canary deployment by configuring the serving endpoint with multiple model versions and specifying a traffic split percentage.
C.Shadow deployment by duplicating requests to a new endpoint without affecting production traffic.
D.Blue-green deployment by creating a new endpoint and switching traffic via DNS.
AnswerB

Databricks Model Serving supports serving multiple model versions on a single endpoint with configurable traffic splits. This enables canary deployments where a small percentage of traffic goes to the new version, allowing gradual rollout and rollback if issues arise. This is the native and recommended strategy for safe updates.

Why this answer

Databricks Model Serving allows you to configure multiple model versions on a single endpoint and set traffic split percentages. This enables canary deployments, where you can gradually shift traffic to a new version while monitoring performance, ensuring high availability and easy rollback. Other strategies either cause downtime or are not natively supported.

Exam trap

The trap here is assuming that classic deployment strategies like blue-green or rolling updates apply directly, but Databricks Model Serving provides built-in traffic splitting for canary deployments.

102
MCQhard

An organization needs to implement a robust CI/CD strategy for their Databricks ML models. Which TWO of the following practices are recommended to ensure reliable model deployment?

A.Perform manual model registration directly from the production workspace notebooks.
B.Use Databricks Repos to synchronize code across development, staging, and production environments.
C.Execute all training runs directly in the production workspace for maximum accuracy.
D.Implement automated model validation tests before promoting a model to the 'Production' stage.
E.Store all raw data directly inside the MLflow model artifact for portability.
AnswerB, D

Databricks Repos allows teams to manage code using Git, which is essential for CI/CD. Synchronizing code across environments ensures that the exact same logic tested in staging is deployed in production, minimizing the risk of 'it works on my machine' errors and facilitating smooth, reproducible deployments.

Why this answer

Implementing automated testing and environment isolation are cornerstones of MLOps. Automated unit and integration tests catch regressions early, while separating development, staging, and production workspaces prevents accidental interference with live systems. Using Git-based workflows via Databricks Repos ensures code traceability, while the Model Registry acts as the gatekeeper for model promotion, ensuring only validated models reach production environments.

Exam trap

Candidates often pick options involving manual model retries or simple notebook exports, failing to recognize that CI/CD requires Git integration and automated gatekeeping via the Model Registry.

103
MCQmedium

What is the best way to handle secrets (like API keys for external feature sources) within a Databricks notebook during model development?

A.Hardcode the keys in a configuration Python script and import it into the notebook.
B.Use the Databricks Secrets API to retrieve credentials at runtime.
C.Store the secrets in an environment variable on the cluster config.
D.Prompt the user to enter the secret manually when the notebook runs.
AnswerB

The Databricks Secrets API allows developers to store sensitive information securely within Databricks and retrieve it programmatically. This keeps credentials out of the codebase entirely, ensuring that only users with the appropriate permissions can access them and that secrets are never logged or stored in version control systems.

Why this answer

Using the Databricks Secrets API is the only secure way to manage credentials. By referencing keys through the `dbutils.secrets.get()` function, developers ensure that sensitive information is never hardcoded or stored in clear text within the version-controlled code. This practice prevents unauthorized access to external systems and is a mandatory security requirement for any enterprise environment dealing with sensitive or paid data APIs.

Exam trap

Candidates frequently suggest using environment variables or configuration files, which are insecure, rather than the Databricks-native Secrets API designed specifically for secure credential management in notebooks.

104
MCQeasy

In the context of Databricks MLOps, what is the primary purpose of a 'Staging' environment in the Model Registry?

A.To act as a long-term backup for model version history.
B.To allow for final validation and testing before production deployment.
C.To automatically convert models into a different serving format.
D.To increase the model's training accuracy.
AnswerB

Staging acts as a testing sandbox where teams can run integration tests or shadow deployments against live data. This ensures that the model is fully vetted for stability and performance in a controlled environment similar to production, mitigating the risk of issues when the model goes live.

Why this answer

The Staging environment serves as a pre-production gate where models undergo final validation, such as integration testing and UAT (User Acceptance Testing), before being promoted to Production. This ensures that only models that meet performance, safety, and operational standards are exposed to live production traffic, reducing the risk of catastrophic failures in the real-world deployment environment.

Exam trap

Test-takers frequently confuse the 'Staging' environment with a live testing zone for end-users, missing its true purpose as an automated pre-production validation gate.

105
Multi-Selecthard

A data science team is preparing to deploy a high-throughput recommendation model using Databricks Model Serving. Which TWO factors must be considered to optimize endpoint latency and resource utilization? (Choose two)

Select 2 answers
A.Configuring the minimum and maximum number of concurrent scaling units based on expected query traffic spikes.
B.Converting all feature lookup tables from Delta Lake format directly into local CSV files inside the serving container.
C.Selecting the appropriate CPU or GPU compute instance type matching the model's computational complexity and memory footprint.
D.Enabling Apache Spark adaptive query execution on the serving endpoint cluster to optimize joins during scoring.
E.Ensuring the MLflow model uses pickle serialization exclusively because it outperforms all other serialization formats.
AnswersA, C

Configuring scaling units correctly allows Databricks Model Serving to automatically scale compute resources up during traffic surges and scale down to zero during idle periods. This balance prevents latency degradation under heavy load while minimizing unnecessary cloud expenditure when request volumes drop.

Why this answer

Selecting appropriate instance types with GPU acceleration and configuring autoscaling bounds based on traffic patterns directly dictate inference latency and operational cost. These considerations are critical in production because underprovisioned endpoints cause timeout errors while overprovisioned endpoints waste cloud computing resources unnecessarily.

Exam trap

Candidates choose data storage options or training parameters instead of focusing on runtime compute configurations and scaling boundaries that directly affect latency and throughput.

106
MCQmedium

A machine learning engineer needs to track model experiments and ensure that all training parameters and metrics are captured in a reproducible way. What is the best practice for using MLflow within Databricks notebooks?

A.Manually log metrics to a shared Excel file stored on DBFS after the run completes.
B.Use MLflow with 'mlflow.start_run()' and 'mlflow.autolog()' to ensure all runs are tracked automatically.
C.Only log the final evaluation metric at the end of the training loop to minimize storage.
D.Print the parameters to the notebook output cell and rely on the notebook's revision history.
AnswerB

Using 'mlflow.start_run()' within a context manager ensures that all tracking information is cleaned up correctly, even if the code fails. 'mlflow.autolog()' reduces boilerplate code by capturing common model parameters and metrics automatically, significantly improving the consistency and efficiency of the experiment tracking process across the organization.

Why this answer

Using 'mlflow.autolog()' or explicit 'mlflow.log_params()' within a context manager ensures that every run is isolated and documented. This practice allows for easy comparison of different hyperparameter configurations, preventing data loss and providing a clear lineage from the raw data to the final model artifact, which is crucial for compliance and debugging in enterprise ML environments.

Exam trap

Candidates often mention manual logging of every parameter individually, missing that 'mlflow.autolog()' is the best practice for capturing comprehensive run data with minimal code overhead.

107
MCQmedium

A team is developing a model to forecast demand. They need to ensure that their feature engineering code is reusable for both training and real-time inference. Which architectural pattern should they adopt?

A.Store preprocessing logic in individual notebook files and import them dynamically.
B.Create a Feature Store table and register feature engineering as a Feature Spec.
C.Perform all feature engineering inside the model's 'predict' method.
D.Write all features to a CSV file and load them during inference.
AnswerB

Feature Specs allow for consistent, reusable feature transformations. By defining the logic in the Feature Store, the team ensures the same preprocessing code is used for both training and online serving. This standardizes the pipeline, reduces engineering effort, and prevents training-serving skew, ensuring that models perform consistently in production.

Why this answer

Defining features within the Databricks Feature Store allows the same code to be used for batch training and real-time lookup. This eliminates the risk of training-serving skew, a major cause of production model degradation. By centralizing feature engineering, teams ensure that the transformations applied during training are identical to those applied during inference, maintaining consistency across the entire development and deployment pipeline.

Exam trap

Candidates often suggest manual code duplication or custom SQL scripts for inference, failing to recognize that the Feature Store is the standard pattern for preventing training-serving skew.

108
MCQhard

A team runs a nightly Databricks job that retrains a demand-forecasting model and registers a new version in the MLflow Model Registry. Compliance requires that the model version used for scoring in production be immutably identified and that any subsequent retraining not silently change what production serves. Which practice best satisfies this requirement?

A.Register each retrained model under a new registered model name that encodes the training date, and update the scoring job to read the newest name each night.
B.Store the model artifact path from the training run in a Delta table and have the scoring job read the path directly, bypassing the Model Registry entirely.
C.Have the scoring job reference the model by its registered model name only, so that the latest version is always used after each nightly run.
D.Assign a stage or alias to the specific model version and have the scoring job reference the model by that stage or alias, updating it only through a controlled promotion step.
AnswerD

Stages and aliases provide an indirection pointer to a specific model version. By promoting only through a controlled step, the scoring job always resolves to the exact version that was validated, and retraining that registers new versions does not alter production until an explicit promotion changes the pointer, satisfying immutability and traceability.

Why this answer

Pinning production to a specific model version through a stage or alias gives an auditable pointer that only changes via an explicit promotion. Referencing the model name alone, inventing new names per run, or reading raw artifact paths all allow production to drift or lose lineage, so they fail the immutability and traceability requirement.

Exam trap

The trap here is believing that referencing the registered model name always serves a fixed artifact, when it actually resolves to whatever version the current stage or alias points to.

109
MCQmedium

Your team is using MLflow Model Registry to manage a model that predicts customer churn. A new version has been registered and passed validation. You need to transition this version to the 'Production' stage and ensure that all downstream scoring jobs automatically use it. What should you do?

A.Keep the model version in 'Staging' and update the downstream jobs to reference the model by stage 'Staging'.
B.Transition the model version to 'Production' and update the downstream jobs to reference the model by stage 'Production'.
C.Archive the current production model and register the new version with a new model name.
D.Transition the model version to 'Production' and update the downstream jobs to reference the model by version number.
AnswerB

Transitioning to 'Production' and having jobs reference the stage ensures that when a new version is promoted, the jobs automatically pick it up. This is the intended workflow for stage-based model deployment in MLflow Model Registry, enabling seamless updates without code changes.

Why this answer

Transitioning the model version to the 'Production' stage and having downstream jobs reference the model by stage 'Production' is the correct approach. This allows you to promote new versions to production without modifying job code, ensuring that all consumers automatically use the latest approved model. It also maintains a clear audit trail and governance.

Exam trap

The trap here is confusing version-based referencing with stage-based referencing; only stage-based referencing enables automatic updates when a new version is promoted.

110
MCQeasy

A team wants to deploy a scikit-learn model to a real-time REST endpoint on Databricks. They have logged the model with MLflow and registered it in Unity Catalog. Which method should they use to create the serving endpoint?

A.Use the Databricks Model Serving UI or REST API to create an endpoint and select the registered model version.
B.Package the model into a Docker image and deploy it to a Kubernetes cluster managed by Databricks.
C.Write a Databricks job that loads the model and exposes it via a Flask app on a driver node.
D.Use MLflow's `mlflow models serve` command on a Databricks cluster to start a local REST server.
AnswerA

Databricks Model Serving integrates directly with Unity Catalog and MLflow. You create an endpoint via the Serving UI or the REST API, specifying the full model name and version. The service automatically provisions the necessary infrastructure and builds a container with the model's dependencies, making this the standard and supported deployment path.

Why this answer

Databricks Model Serving is the managed solution for deploying registered models as REST endpoints. It handles infrastructure, scaling, and dependency management automatically. Using the UI or REST API to create an endpoint from a Unity Catalog model version is the correct and supported approach, unlike manual containerization or local serving.

Exam trap

The trap here is confusing local MLflow serving commands or custom Flask apps with the managed Databricks Model Serving feature.

111
MCQhard

A team notices that their model performance is significantly lower in production than in training. They suspect 'data drift' in the feature inputs. Which Databricks capability should be used to monitor this?

A.Manually querying the Delta table after every batch inference.
B.Using Databricks Lakehouse Monitoring to track inference data distributions.
C.Enabling 'Audit Logs' to track individual inference request parameters.
D.Increasing the number of features in the model to cover more data.
AnswerB

Lakehouse Monitoring automates the process of comparing current inference data distributions against training data snapshots. It identifies statistically significant drift, providing the necessary visibility for proactive model maintenance. This tool is standard for maintaining model performance in production, ensuring that decay is detected and addressed promptly.

Why this answer

Databricks Lakehouse Monitoring provides automated insights into data quality and drift by comparing production inference data against training baselines. This is essential for detecting the performance decay caused by concept or data drift. By identifying these issues early, teams can trigger automated retraining pipelines, maintaining model accuracy over time and ensuring the business value of the deployed model is sustained.

Exam trap

Candidates often suggest retraining the model immediately or checking training metrics, failing to recognize that production drift requires dedicated monitoring tools.

112
MCQeasy

A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the one with the lowest validation RMSE. Which MLflow UI feature allows them to sort and filter runs by a specific metric?

A.The Notebook's `mlflow.search_runs` output.
B.The Model Registry's version list.
C.The Artifacts tab within a run.
D.The Runs table with column sorting and metric filters.
AnswerD

The MLflow UI's Runs table displays runs with columns for parameters, metrics, and tags. Users can sort by a metric column and apply filters to narrow down runs. This is the standard way to compare runs and find the one with the lowest RMSE. It provides a quick visual comparison.

Why this answer

The MLflow UI's Runs table is designed for comparing runs. It allows sorting by metrics such as RMSE and filtering runs based on metric values or tags. This makes it straightforward to identify the best run.

Other tabs or the Model Registry serve different purposes.

Exam trap

The trap here is confusing the Model Registry or notebook output with the interactive Runs table in the MLflow UI.

113
MCQhard

A data scientist is using MLflow to log a model that includes a custom preprocessing step implemented in Python. They want to ensure that the preprocessing logic is packaged with the model so that it can be served consistently. Which MLflow model flavor should they use?

A.`mlflow.tensorflow`
B.`mlflow.pytorch`
C.`mlflow.pyfunc`
D.`mlflow.sklearn`
AnswerC

The `mlflow.pyfunc` flavor allows packaging any Python model with custom preprocessing and prediction logic. By subclassing `PythonModel` and implementing `predict`, the data scientist can include preprocessing steps directly in the model artifact. This ensures the entire logic is self-contained and reproducible during serving. It is the most flexible flavor for custom code.

Why this answer

The `mlflow.pyfunc` flavor is designed for custom Python models and allows embedding arbitrary preprocessing and postprocessing logic. By creating a `PythonModel` subclass, the data scientist can package the entire inference pipeline, ensuring consistency between training and serving. Other flavors are library-specific and may not capture custom Python steps outside their frameworks.

Thus, `mlflow.pyfunc` is the correct choice.

Exam trap

The trap here is assuming that a library-specific flavor like `mlflow.sklearn` will automatically package custom Python preprocessing, when it may only save the model object and require external code.

114
MCQmedium

A data scientist has deployed a model to Databricks Model Serving and wants to monitor its performance over time. They need to track prediction drift and data quality issues. Which Databricks feature should they use to automatically capture inference logs and compute metrics?

A.Set up a Databricks job to periodically query the endpoint.
B.Use the model's signature to validate incoming data.
C.Configure MLflow tracking to log model predictions.
D.Enable inference tables on the serving endpoint.
AnswerD

Inference tables automatically capture request and response data from a Model Serving endpoint and store them in a Delta table. This enables monitoring of prediction drift, data quality, and model performance by analyzing the logged data. It is the native Databricks feature designed for this purpose, integrating with Lakehouse Monitoring.

Why this answer

Inference tables are a Databricks feature that automatically logs request and response payloads from a Model Serving endpoint into a Delta table. This data can then be used with Lakehouse Monitoring to track prediction drift, data quality, and model performance. Other options do not provide automatic, scalable logging of production inference data.

Exam trap

The trap here is assuming that MLflow tracking or manual queries can replace inference tables for production monitoring, but they lack automatic capture and integration with monitoring tools.

115
MCQhard

You are implementing a CI/CD pipeline for a model. You want to automate unit testing of the model's inference performance. Which approach is best suited for Databricks?

A.Manually inspect performance metrics in MLflow UI before every deployment.
B.Trigger a notebook to run inference on a validation set and assert metrics.
C.Deploy the model to staging and wait for user feedback.
D.Rely on the Model Registry to automatically validate model performance.
AnswerB

Automating validation via notebooks allows for programmatic assertion of metrics like accuracy or latency. This creates a reliable quality gate that runs as part of the deployment pipeline. If the model fails the assertions, the pipeline halts, preventing the promotion of a sub-par model to production.

Why this answer

Using a dedicated job workflow that triggers a notebook to run inference on a hold-out test dataset is the standard Databricks approach. By comparing metrics against defined thresholds before deployment, you ensure quality. This automated gate prevents poor-performing models from reaching production, which is a critical step in a mature MLOps lifecycle to maintain system stability and model accuracy.

Exam trap

Candidates often suggest using external CI/CD tools like Jenkins or GitHub Actions to perform model testing, forgetting that Databricks jobs can natively run notebooks to validate inference performance.

116
MCQmedium

A data scientist is using Databricks AutoML to solve a classification problem. After the run completes, they want to modify the feature engineering logic for the best-performing model. Which artifact should they retrieve from the AutoML run?

A.The raw MLflow model artifact stored in S3/DBFS.
B.The AutoML generated 'best trial' notebook.
C.The MLflow experiment metrics file in JSON format.
D.The database schema definition for the input features.
AnswerB

The 'best trial' notebook contains all the preprocessing and training code that produced the top-performing model. By cloning this notebook, the data scientist gains full control over the feature engineering pipeline, allowing them to iterate on transformations and custom logic that AutoML initially discovered or implemented automatically.

Why this answer

Databricks AutoML generates a 'data exploration' notebook and a 'best trial' notebook for every experiment. The notebook contains the full Python code used for feature engineering, model training, and hyperparameter tuning. Accessing this notebook allows data scientists to treat AutoML results as a baseline, enabling them to refine feature engineering steps or adjust model parameters for further performance improvements.

Exam trap

Candidates often try to manually inspect the model binary or logs, forgetting that AutoML provides a full, editable notebook that contains the exact feature engineering logic used.

117
Multi-Selecthard

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure the endpoint can handle sudden spikes in traffic without downtime. The model has a large memory footprint and takes several seconds to load. Which TWO configurations should the engineer implement to achieve this? (Choose two.)

Select 2 answers
A.Set a low maximum concurrency per instance to force more instances to be created.
B.Set the minimum provisioned concurrency to a value greater than zero to keep instances warm.
C.Configure the endpoint to use a larger workload size to accommodate the model's memory footprint.
D.Use a smaller model version to reduce load time.
E.Enable scale-to-zero to reduce costs during periods of no traffic.
AnswersB, C

Setting a minimum provisioned concurrency ensures that a specified number of model instances are always running, ready to handle requests. This eliminates cold-start latency during traffic spikes, as instances are pre-loaded with the model. For a model with a large memory footprint and slow load time, this is crucial to avoid downtime and maintain performance.

Why this answer

To handle traffic spikes without downtime for a model with a large memory footprint and slow load time, the engineer should keep instances warm by setting a minimum provisioned concurrency greater than zero, and ensure sufficient resources by selecting a larger workload size. These two configurations work together to provide immediate capacity and prevent cold starts, ensuring the endpoint remains responsive under sudden load.

Exam trap

The trap here is focusing solely on scaling out quickly while overlooking the need to keep warm instances and provide adequate memory, which are essential for models with slow load times and large footprints.

118
MCQmedium

Which THREE of the following are key responsibilities of an MLOps engineer when managing a model lifecycle on Databricks?

A.Automating the model registration and promotion process in CI/CD.
B.Writing the core machine learning algorithms for every project.
C.Monitoring the production health and data drift of deployed models.
D.Managing access control and governance of model artifacts.
E.Manually testing every model in the production environment.
AnswerA, C, D

Automation is at the heart of MLOps. By scripting the model registry interactions within CI/CD pipelines, engineers ensure that models move through stages consistently, preventing manual errors and ensuring that the promotion process is documented and repeatable across the organization's different deployment environments.

Why this answer

MLOps engineers act as the bridge between data science and production operations. Their core duties involve automating the lifecycle, ensuring that data quality is monitored, and maintaining the infrastructure security and governance required to serve models reliably. These activities ensure that machine learning remains a stable and repeatable business function rather than an ad-hoc, error-prone research endeavor.

Exam trap

Candidates frequently select ad-hoc data science tasks like hyperparameter tuning or feature engineering as primary MLOps responsibilities, ignoring governance, CI/CD promotion automation, and production monitoring.

119
Multi-Selectmedium

A financial institution is using Databricks to build and deploy a credit risk model. The model must comply with regulations that require full auditability of the model's lineage, including data sources, transformations, and training parameters. The team uses MLflow for tracking and the Feature Store for feature management. Which TWO of the following practices are essential to meet the auditability requirements? (Choose two.)

Select 2 answers
A.Use Databricks Feature Store to create feature tables and log the training set with the model, including feature lookup keys.
B.Log all input data versions and feature table versions used in training as MLflow tags or artifacts.
C.Enable MLflow autologging for all supported frameworks to capture parameters and metrics automatically.
D.Manually document the data sources and transformations in a separate spreadsheet for each model version.
E.Store all training code in a Git repository and commit after each run.
AnswersA, B

The Feature Store allows you to create feature tables with versioning and lineage. When you train a model using features from the Feature Store and log the model with mlflow.pyfunc.log_model, it automatically records the feature lookup keys and the feature table versions. This creates a direct link between the model and the features used, satisfying audit requirements for feature lineage.

Why this answer

To meet auditability requirements, it is essential to automatically capture data and feature lineage. Logging data versions and feature table versions in MLflow provides a record of the exact inputs used. Using the Feature Store to log the training set with the model ensures that feature lineage is preserved.

Together, these practices create a transparent and verifiable chain from data to model, which is necessary for regulatory compliance.

Exam trap

The trap here is thinking that code versioning or autologging alone suffices for auditability, when the critical missing piece is data and feature lineage.

120
MCQmedium

When deploying a model to Databricks Model Serving, you notice that inference latency is higher than expected. Which diagnostic approach is most effective for identifying the bottleneck?

A.Re-train the model with a smaller dataset to see if it improves performance.
B.Check the built-in request metrics and logs for the serving endpoint to analyze duration distribution.
C.Increase the number of instances in the model serving endpoint without checking logs.
D.Restart the Databricks workspace to clear cached inference results.
AnswerB

The built-in monitoring tools provide granular data on request duration and system performance. Analyzing this distribution allows you to identify if the latency is systematic or limited to specific types of requests, enabling you to pinpoint if the bottleneck lies in compute resources, model execution, or network overhead.

Why this answer

Monitoring tools provided by Databricks, such as the built-in request metrics and logs, are essential for identifying latency bottlenecks. By examining request volume, processing duration, and resource utilization, you can determine if the latency is due to model complexity, infrastructure constraints, or external data dependencies. This allows for data-driven optimization, such as choosing a larger workload size, optimizing the model architecture, or implementing caching for frequently accessed data inputs.

Exam trap

Candidates often look to external APM tools or rewrite model code immediately, overlooking the built-in request metrics and logs readily available directly in the Databricks serving interface.

121
MCQmedium

A data scientist is developing a scikit-learn model on Databricks and wants to log the model artifact to MLflow so that it can later be deployed for online inference. They call mlflow.sklearn.log_model() without providing a signature. What is the primary consequence of omitting the model signature?

A.MLflow will not log the model at all because a signature is mandatory for scikit-learn models.
B.The model will be logged, but MLflow Model Serving will reject requests because it cannot validate the input schema.
C.The model will be logged, but MLflow cannot infer the input and output schema for deployment, potentially causing issues in serving or scoring.
D.MLflow will automatically generate a signature by inspecting the training data, so omitting it has no effect.
AnswerC

Without a signature, MLflow lacks a formal description of the model's expected input and output types. This can cause problems when deploying to Model Serving or when using the model in a batch scoring job that expects schema validation. The model artifact itself is still usable, but the missing schema may lead to integration issues or ambiguous behavior.

Why this answer

Providing a signature when logging a model captures the expected input and output schema, which is essential for reliable deployment and scoring. Without it, MLflow cannot validate incoming data against the model's expectations, potentially leading to errors in serving or batch inference. The model still logs successfully, but the lack of schema metadata can cause integration problems.

Exam trap

The trap here is assuming that a model signature is required for logging or that its absence prevents serving, when in fact it is optional metadata that affects schema validation and deployment robustness.

122
MCQmedium

When evaluating a classification model on Databricks, a team needs to generate a custom performance report that is not natively provided by MLflow. What is the recommended strategy to ensure this report is persisted and associated with the training run?

A.Print the report to the notebook output and rely on the notebook's command history.
B.Save the report to a local temporary file and use MLflow.log_artifact to upload it to the run.
C.Store the report in an external database and manually link it via a comment in the run.
D.Re-generate the report every time the model is loaded for inference.
AnswerB

The 'log_artifact' method is the standard way to attach any file, including custom reports, images, or plots, to an MLflow run. This provides a robust and centralized way to keep evaluation results alongside the model, enabling comprehensive documentation of the model lifecycle directly within the Databricks MLflow tracking environment.

Why this answer

Logging custom plots and reports as artifacts is the recommended way to enrich MLflow run metadata. By saving visualizations or summary statistics as files (e.g., HTML, PNG, or JSON) and using 'log_artifact', the team ensures that non-standard evaluation metrics are permanently stored. This practice allows stakeholders to review specific model performance indicators without needing to re-run the training code, which is vital for compliance and post-training analysis.

Exam trap

Candidates often try to use direct print statements or UI screenshot functions, forgetting that custom programmatic reports must be saved as local files first before being uploaded via log_artifact.

123
Multi-Selecthard

You are building a CI/CD pipeline that must promote an MLflow model version from Staging to Production in Databricks only after automated validation. The pipeline runs in a service principal context. Which two actions are required to implement this safely and repeatably? (Choose two.)

Select 2 answers
A.Grant the service principal CAN_MANAGE_PRODUCTION (or CAN_MANAGE) on the registered model so it can transition versions into Production.
B.Configure the pipeline to call the MLflow Model Registry transition-stage API after validation succeeds, using the model name and version.
C.Enable automatic model version archiving so that older Production versions are removed before the new version is promoted.
D.Require the service principal to use a personal access token tied to a human user so that audit logs show a named approver.
E.Store the model artifacts in a Unity Catalog volume and reference the volume path when registering the model version.
AnswersA, B

Stage transitions in MLflow Model Registry are governed by model-level permissions. A service principal running the pipeline needs at least CAN_MANAGE_PRODUCTION or CAN_MANAGE to move a version into Production. Without this grant, the API call to transition the stage will fail with a permission error, making the pipeline non-repeatable. This is a required, concrete step for automated promotion.

Why this answer

Safe automated promotion needs two things: permission for the automation identity to change stages, and a programmatic call to perform the transition after validation. Granting the service principal CAN_MANAGE_PRODUCTION or CAN_MANAGE satisfies the first, and invoking the transition-stage API after validation satisfies the second. Archiving policies, artifact storage changes, and human tokens are neither required nor advisable for this pipeline.

Exam trap

The trap here is focusing on artifact storage or archiving policies, when the required elements are the service principal's model permission and the programmatic stage transition call.

124
MCQeasy

A data scientist is using MLflow to log a model trained with scikit-learn. They want to ensure that the model can be loaded later for batch inference using `mlflow.pyfunc.load_model`. Which condition must be met for the model to be loadable as a PyFunc model?

A.The model must be saved in the ONNX format to be compatible with PyFunc.
B.The model must be logged with a signature that defines the input schema; otherwise, PyFunc loading will fail.
C.The model must be logged with `mlflow.sklearn.log_model`, which automatically adds the `python_function` flavor.
D.The model must be registered in the MLflow Model Registry before it can be loaded as a PyFunc model.
AnswerC

When logging a scikit-learn model with `mlflow.sklearn.log_model`, MLflow automatically includes the `python_function` flavor. This makes the model loadable via `mlflow.pyfunc.load_model`, which provides a generic interface for inference. The Python function flavor wraps the native model, allowing it to be used in environments where the original library may not be available or for consistent serving.

Why this answer

To load a scikit-learn model as a PyFunc model, it must be logged with `mlflow.sklearn.log_model`, which automatically includes the `python_function` flavor. This flavor provides a consistent inference API across different model types. Registration is not required, nor is a signature or ONNX format.

The key is that the logged model contains the PyFunc flavor, which `mlflow.pyfunc.load_model` uses to load and serve the model.

Exam trap

The trap here is thinking that Model Registry registration is required to load a model as PyFunc, when actually the PyFunc flavor is automatically added during logging.

125
MCQhard

An ML engineer is deploying a model to a Databricks Model Serving endpoint. The model's inference function logs predictions to a Delta table for monitoring. During testing, they notice that the logging adds significant latency. They need to reduce the impact on inference latency. Which approach should they take?

A.Enable auto-scaling to handle the additional load from logging.
B.Reduce the frequency of logging by sampling predictions.
C.Move the logging to an asynchronous background process.
D.Increase the workload size of the serving endpoint.
AnswerC

Asynchronous logging decouples the logging operation from the inference path. The model returns predictions immediately, and logging occurs in the background. This significantly reduces latency because the response is not delayed by I/O operations. Databricks Model Serving supports asynchronous logging via custom code or by using the inference table feature, which logs asynchronously.

Why this answer

Asynchronous logging moves the logging operation outside the critical path of inference, allowing predictions to be returned without waiting for the log write to complete. This is the most effective way to reduce latency caused by logging in a serving endpoint. Databricks Model Serving supports asynchronous logging patterns, such as using the inference table or custom async code.

Exam trap

The trap here is assuming that scaling resources or reducing logging frequency will eliminate the latency, when the real issue is synchronous blocking I/O.

126
MCQmedium

Refer to the exhibit. A data scientist is logging their model training process. Which statement accurately describes the storage location of the artifacts referenced in the code snippet?

A.All items are stored in the same local directory on the driver node.
B.The model artifact created by log_model is stored in the MLflow tracking store's internal directory structure.
C.The log_artifact command uploads the weights directly to the Model Registry.
D.The log_model function is equivalent to copying the file directly to DBFS.
AnswerB

The 'log_model' function creates a standardized structure containing the model binary, a MLmodel file, and a conda.yaml file. This allows MLflow to maintain environment parity during deployment. The tracking store automatically manages these artifacts, abstracting the underlying storage layer, which ensures consistency across different stages of the ML lifecycle.

Why this answer

MLflow handles artifacts differently based on the log function used. While 'log_artifact' takes a explicit path (often DBFS or local filesystem), 'log_model' packages the model with its metadata and dependencies into a standard MLflow directory structure. Understanding this distinction is essential for production deployments, as model registry and deployment tools rely on the specific internal structure generated by the 'log_model' function, not just raw file paths.

Exam trap

Candidates often confuse log_model with log_artifact, assuming that model artifacts are dumped into a generic user-specified folder rather than structured automatically within MLflow's internal tracking store directory.

127
Multi-Selecthard

You are implementing a CI/CD pipeline for ML models on Databricks. The pipeline must automatically retrain, validate, and promote models to Production in MLflow Model Registry. Which TWO practices are essential for maintaining reproducibility and governance in this automated workflow? (Choose two.)

Select 2 answers
A.Configure the pipeline to automatically transition any model version with accuracy above a fixed threshold directly to Production without human review.
B.Disable model versioning in the registry and instead use run IDs as the sole identifier for production models.
C.Use MLflow Model Registry stage transitions with recorded descriptions and tags that identify the pipeline run and approver for each promotion.
D.Log the Git commit SHA and environment specification with each MLflow run so the exact code and dependencies can be reproduced.
E.Store model artifacts in a separate cloud storage bucket outside the MLflow tracking server to avoid coupling with the registry.
AnswersC, D

Recording descriptions and tags on stage transitions provides an audit trail linking each promotion to a specific pipeline run and approver. This supports governance by making it possible to trace who or what promoted a model and why, which is essential for compliance and incident investigation.

Why this answer

Reproducibility and governance in automated pipelines require capturing the exact code and environment for each run, and maintaining an auditable record of stage transitions with descriptions and tags. Automatic promotion without review and decoupling artifacts or disabling versioning all weaken traceability and control, making them unsuitable for a governed CI/CD workflow.

Exam trap

The trap here is equating automated promotion with governance, when governance actually requires traceability of code, environment, and promotion decisions.

128
MCQmedium

A team wants to track model performance over time to detect drift without writing custom monitoring infrastructure. What is the most efficient Databricks tool for this?

A.MLflow Tracking APIs.
B.Databricks Model Monitoring (DMM).
C.Delta Lake Change Data Feed.
D.Databricks SQL Dashboards.
AnswerB

DMM is the native solution for monitoring production models. It automates drift detection, schema validation, and performance tracking, providing ready-made dashboards and alerts. It eliminates the need to build custom monitoring infrastructure, making it the most efficient choice for teams wanting to maintain high-quality models in production.

Why this answer

Databricks Model Monitoring (DMM) provides built-in capabilities to track model performance, data drift, and input quality. By automating these tasks, teams avoid the technical debt of building and maintaining custom monitoring solutions. DMM integrates seamlessly with the Model Registry, allowing for automated alerts and reports that keep stakeholders informed about model health in production environments.

Exam trap

Test-takers might suggest writing custom Spark jobs or MLflow logging scripts, unaware that Databricks Model Monitoring provides native, out-of-the-box tracking.

129
Multi-Selecthard

A machine learning engineer is preparing a scikit-learn model for batch scoring with MLflow on Databricks. The team wants the logged model to carry a reproducible environment and a machine-readable description of the input and output schema so downstream consumers can validate requests. Which TWO actions should the engineer take when logging the model with `mlflow.sklearn.log_model`? (Choose two.)

Select 2 answers
A.Set `pip_requirements` to a pinned requirements file or list of exact package versions.
B.Set `registered_model_name` to publish the model immediately to the Registry.
C.Pass `artifacts` pointing to the training notebook for reference.
D.Pass a `signature` built from a sample input DataFrame using `mlflow.models.infer_signature`.
E.Set `serialization_format` to `cloudpickle` to embed the environment.
AnswersA, D

Specifying `pip_requirements` captures the exact library versions needed to recreate the training environment. MLflow stores this in the model's conda and requirements metadata, so a later restore recreates the same dependency set. This is how the team achieves a reproducible environment rather than relying on whatever happens to be installed on the scoring cluster.

Why this answer

Reproducibility and schema validation are two distinct metadata concerns addressed by separate parameters. Pinning dependencies through `pip_requirements` records the exact package environment so the model can be restored consistently later. Building a `signature` with `infer_signature` records column names, types, and the output schema so scoring endpoints and the Registry can enforce input contracts.

Together they make the logged model self-describing and portable across clusters.

Exam trap

The trap here is conflating registration and serialization settings with environment pinning and schema capture, which are handled by different log_model arguments.

130
MCQhard

Refer to the exhibit. A data scientist attempted to register a new model version but received the error shown in the exhibit. Which step should be taken to resolve this issue?

A.Upgrade the user's Databricks entitlement to 'Workspace Admin'.
B.Update the MLflow experiment tracking URI to point to an external database.
C.Ask an administrator to grant 'CAN_MANAGE' or 'CAN_EDIT' permissions on the model in the MLflow UI.
D.Re-run the training notebook using the cluster owner's credentials.
AnswerC

Databricks implements model-specific permissions. To register or modify model versions, the user must have the appropriate permission level assigned to that specific model object. This follows standard security best practices by limiting access to authorized users without granting excessive, unnecessary privileges across the entire workspace or environment.

Why this answer

The error indicates a lack of necessary permissions to perform operations on the MLflow Model Registry in the Databricks workspace. Access control in Databricks is granular; even if a user can write code, they need specific registry-level permissions to promote or register models. This is a common MLOps governance issue where security policies must be correctly balanced with developer autonomy.

Exam trap

Candidates often suggest re-training the model or checking the code, overlooking that this is an IAM/RBAC issue within the Databricks Workspace specific to the Model Registry.

131
Multi-Selecthard

A machine learning engineer is preparing to deploy a model to production using MLflow Model Registry. They want to ensure that the model can be easily served and that its dependencies are correctly captured. Which TWO actions should they take when logging the model to guarantee that the serving environment can recreate the necessary Python environment? (Choose two.)

Select 2 answers
A.Log the model using `mlflow.sklearn.log_model()` with the `pip_requirements` parameter set to a list of pip requirement strings.
B.Log the model using `mlflow.sklearn.log_model()` with the `code_path` parameter to include custom transformation code, which also captures all dependencies.
C.Log the model using `mlflow.sklearn.log_model()` with the `signature` parameter to infer the input schema and automatically generate the environment.
D.Log the model using `mlflow.sklearn.log_model()` with the `conda_env` parameter specifying a Conda environment YAML file that includes all dependencies.
E.Log the model using `mlflow.sklearn.log_model()` with the `registered_model_name` parameter to register the model, which automatically captures the environment.
AnswersA, D

The `pip_requirements` parameter allows you to specify a list of pip requirements, which MLflow will use to create a `requirements.txt` file. This is an alternative to Conda and is particularly useful in environments where Conda is not available. It ensures that the serving environment installs the correct packages, though it may not capture non-Python dependencies as comprehensively as Conda.

Why this answer

To ensure the serving environment can recreate dependencies, you must explicitly specify them when logging the model. Using the `conda_env` parameter with a Conda YAML file or the `pip_requirements` parameter with a list of pip requirements are the two supported ways to capture dependencies. Both methods result in MLflow saving the necessary environment files with the model, which are then used during serving to install the correct packages.

Exam trap

The trap here is assuming that registering the model or providing a signature automatically captures dependencies, when in fact MLflow requires explicit environment specification for reproducibility.

132
MCQhard

A machine learning engineer is building a feature engineering pipeline in Databricks using Feature Store. They need to ensure that the same feature computation logic is used for both training and batch scoring, and that features are automatically refreshed. Which approach should they take?

A.Create a Feature Store table using a PySpark UDF that reads from a Delta table and schedule a job to refresh it.
B.Write a Python function that computes features and call it manually in both training and scoring notebooks.
C.Use MLflow to log the feature engineering code as an artifact and manually apply it during scoring.
D.Define a feature computation function using the Feature Store client, register it, and create a Feature Store table with a scheduled refresh.
AnswerD

Defining a feature computation function with the Feature Store client and registering it ensures that the same logic is used for training and scoring. Creating a Feature Store table with a scheduled refresh automatically updates features. This approach provides consistency, versioning, and lineage. It directly satisfies the need for identical feature computation and automatic refresh, making it the correct solution for the engineer's pipeline.

Why this answer

Defining a feature computation function and registering it with Feature Store ensures that training and scoring use identical logic. Creating a Feature Store table with scheduled refresh automates feature updates. This combination provides consistency, reproducibility, and freshness.

The other options rely on manual steps or lack the necessary integration with Feature Store, so they do not guarantee consistent feature computation and automatic refresh.

Exam trap

The trap here is assuming that any scheduled job or MLflow artifact can replace Feature Store's feature function and materialization for consistent training and scoring.

133
MCQeasy

A data scientist has trained a model and registered it in Unity Catalog. They now need to deploy it for real-time inference with automatic scaling and a REST API endpoint. Which Databricks feature should they use?

A.Unity Catalog functions
B.Databricks Model Serving
C.Databricks Jobs with a serving cluster
D.MLflow Model Registry webhooks
AnswerB

Databricks Model Serving provides a fully managed, serverless solution to deploy models as REST API endpoints with automatic scaling based on traffic. It integrates with Unity Catalog for model governance and supports real-time inference. This is the standard feature for deploying models for real-time serving in Databricks, offering built-in monitoring and scaling without managing infrastructure.

Why this answer

Databricks Model Serving is the purpose-built feature for deploying models as real-time REST endpoints with automatic scaling. It handles infrastructure, scaling, and monitoring, allowing data scientists to focus on model development. Other options are either for automation, batch processing, or data governance, and do not provide the required real-time serving capabilities.

Exam trap

The trap here is confusing model registry event triggers with actual serving infrastructure, or assuming that batch jobs can handle real-time requests.

134
MCQmedium

Which THREE of the following are considered best practices for handling data preprocessing in a Databricks ML pipeline to prevent data leakage?

A.Fit your scaler/transformer exclusively on the training portion of the data.
B.Calculate global mean imputation values using the entire dataset before splitting.
C.Use Spark ML Pipeline objects to chain preprocessing and model training steps.
D.Ensure that test data is strictly isolated from the preprocessing pipeline.
E.Perform feature selection using the entire dataset to maximize model power.
AnswerA, C, D

Fitting on only the training data ensures that the model does not incorporate information from the validation or test sets. This is the most effective way to prevent data leakage during preprocessing, ensuring that the model's evaluation metrics are realistic and reflect its actual performance on unseen, future production data.

Why this answer

Preventing data leakage is essential for valid model performance. The key is to calculate statistics (like means or scaling factors) strictly on the training set and apply them to the validation/test sets. Using Spark ML transformers encapsulates this logic, ensuring that the transformation process is consistent and prevents the model from 'seeing' information from the test set during the training phase, leading to accurate performance estimation.

Exam trap

Candidates often perform scaling or imputation on the entire dataset before splitting, which is the most common cause of data leakage and leads to overly optimistic model performance.

135
Multi-Selecthard

You are implementing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically retrain the model when new data arrives, validate its performance, and promote it to production if it meets quality thresholds. Which TWO of the following steps are essential to include in the pipeline to ensure safe and automated deployment? (Choose two.)

Select 2 answers
A.Implement automated tests that evaluate the model on a holdout dataset and compare metrics against a baseline before promotion.
B.Manually review the model's performance metrics in a notebook before allowing the pipeline to proceed.
C.Store the model artifacts in a Git repository alongside the code to ensure versioning.
D.Configure the pipeline to directly overwrite the production model endpoint with the newly trained model without validation to reduce latency.
E.Use MLflow Model Registry to register the model and transition it to 'Staging' for validation, then to 'Production' after passing tests.
AnswersA, E

Automated tests on a holdout dataset verify that the model meets performance thresholds before promotion. Comparing against a baseline ensures that the new model does not regress. This step is essential for safe automation, as it gates deployment on objective criteria and prevents degradation in production.

Why this answer

The essential steps for an automated CI/CD pipeline include registering the model in MLflow Model Registry with stage transitions and implementing automated tests on a holdout dataset to validate performance against a baseline. These steps ensure that only models meeting quality thresholds are promoted, enabling safe automation. Other steps like manual review or storing artifacts in Git are not suitable for automated pipelines.

Exam trap

The trap here is thinking that manual review or direct overwrite can be part of an automated pipeline, but automation requires objective, reproducible validation steps without human intervention.

136
Multi-Selectmedium

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle traffic spikes while minimizing costs during idle periods. The engineer considers enabling scale-to-zero and configuring autoscaling. Which TWO statements about these features are correct? (Choose two.)

Select 2 answers
A.Scale-to-zero and autoscaling cannot be enabled simultaneously on the same endpoint.
B.Scale-to-zero reduces replicas to zero after a period of inactivity, which can lead to cold-start latency on the next request.
C.Autoscaling scales based on CPU utilization of the serving containers, not on request concurrency.
D.Scale-to-zero is only available for GPU workload types, not for CPU workload types.
E.Autoscaling automatically adjusts the number of replicas based on incoming request load, up to a maximum configured limit.
AnswersB, E

Scale-to-zero is designed to save costs by scaling down to zero replicas when there is no traffic. When a request arrives after idle time, the endpoint must scale up, causing a cold start and increased latency for that request. This is a fundamental trade-off between cost and latency.

Why this answer

Scale-to-zero reduces replicas to zero during inactivity, causing cold-start latency on the next request. Autoscaling adjusts replicas based on load, up to a maximum. These features can be used together, with scale-to-zero effectively setting the minimum replicas to zero.

Autoscaling uses request concurrency as the primary scaling metric, not CPU utilization.

Exam trap

The trap here is thinking that scale-to-zero and autoscaling are mutually exclusive or that autoscaling scales on CPU, when it actually scales on request concurrency.

137
MCQhard

Refer to the exhibit. A developer wants to ensure the Random Forest model can be used for automated inference at scale. Based on the provided code, what is missing to enable the model to support the 'predict' method within the Databricks Model Serving environment?

A.The model must be registered to the model registry before logging.
B.The model.fit(X_train, y_train) method must be called before log_model.
C.The model must be wrapped in an mlflow.pyfunc.PythonModel class.
D.The code must explicitly define the input signature in log_model.
AnswerB

Scikit-learn models must be trained via the .fit() method before they contain the internal state necessary for inference. Logging an unfitted instance saves an empty shell, which cannot perform predictions. The model must be trained on representative data to ensure the internal weights are populated correctly.

Why this answer

The provided code logs an unfitted model instance. MLflow cannot serialize a model's 'predict' capabilities until the model has actually been trained (fitted) on a dataset. Attempting to deploy or run an unfitted model results in an error, as there is no learned logic to execute.

The model must be fitted with data prior to being passed to log_model.

Exam trap

Candidates often focus on the model type (e.g., Random Forest) or the logging library, missing the fundamental requirement that a model must be trained (fitted) to be serializable.

138
Multi-Selecthard

A fraud-detection model is served via a Databricks Model Serving endpoint. Compliance requires that every prediction request be traceable to the exact model artifact that produced it and that the model's inputs be auditable for drift analysis. Which two actions should you take to satisfy these requirements? (Choose two.)

Select 2 answers
A.Set the endpoint's scale-to-zero behavior to disabled so logs are never lost during cold starts.
B.Configure the endpoint to use a custom container image that writes each request to an external S3 bucket.
C.Enable inference tables on the endpoint so requests and responses are automatically logged to a Delta table.
D.Tag the registered model version in Unity Catalog with the deployment timestamp and serving endpoint name.
E.Enable automatic model version capture on the endpoint so the served model version is recorded with request logs.
AnswersC, E

Inference tables capture the payloads sent to and returned from a Model Serving endpoint and persist them to a Unity Catalog Delta table. This gives an auditable record of the model inputs for drift analysis and links each request to the served model version, satisfying both traceability and auditability requirements without custom logging code.

Why this answer

Inference tables persist endpoint request and response payloads to a governed Delta table, and enabling model version capture records which served model version produced each prediction. Together they provide the per-request audit trail and the artifact-level traceability the compliance scenario demands, without custom containers or manual tagging that cannot capture payloads.

Exam trap

The trap here is assuming that tagging the model version or keeping replicas warm is enough for auditability, when only payload logging and version capture actually record each prediction.

139
MCQhard

An ML engineering team maintains a feature table in Databricks Feature Store that is populated by a nightly batch job. The same features feed both an offline training pipeline and a Model Serving endpoint. After a schema change adds two columns to the source Delta table, the endpoint begins returning errors during online lookup. The team confirms the offline training pipeline still works. Which change most likely restores the endpoint?

A.Publish the updated feature table to the online store so the online table schema matches the offline table, then redeploy the model that consumes those features.
B.Add the new columns to the model signature and register a new model version, leaving the online store untouched.
C.Recreate the serving endpoint with a larger workload size so the added columns can be cached in memory during lookups.
D.Switch the endpoint to read features directly from the offline Delta table instead of the online store.
AnswerA

The online store is a separate materialized copy optimized for low-latency lookup, and it does not automatically inherit schema changes made to the offline Delta table. Publishing again refreshes the online table to include the new columns so the serving lookup matches what the model expects. Redeploying ensures the model signature aligns with the updated feature schema.

Why this answer

Feature Store maintains two representations: the offline Delta table for training and an online table for low-latency serving. Schema changes flow to the offline table automatically but must be republished to refresh the online store. Because the endpoint queries the online store, its stale schema causes lookup errors even though training still succeeds.

Republishing and redeploying restores alignment.

Exam trap

The trap here is assuming the online store automatically tracks schema changes applied to the offline feature table.

140
MCQeasy

An ML engineer needs to deploy a model to Databricks Model Serving that requires a specific version of a Python library. The library is available on PyPI. Where should the engineer specify this dependency?

A.In the Databricks cluster's init script.
B.In the model's conda.yaml file.
C.In the endpoint configuration's environment variables.
D.In the model's MLflow signature.
AnswerB

The conda.yaml file is part of the MLflow model artifact and specifies the environment dependencies, including Python packages from PyPI. When deploying to Databricks Model Serving, the serving infrastructure uses this file to build the environment. Therefore, the engineer should add the library and its version to the pip section of conda.yaml.

Why this answer

For a model deployed to Databricks Model Serving, Python dependencies are specified in the model's conda.yaml file. This file is part of the MLflow model artifact and is used by the serving infrastructure to create the environment. Adding the required library and version to the pip section ensures it is installed and available during inference.

Exam trap

The trap here is assuming that dependencies can be set via environment variables or init scripts, when they must be declared in the model's conda.yaml.

141
MCQmedium

An ML engineer wants to ensure that their model training pipeline is robust against data quality issues. Which approach, if integrated into the pipeline, most effectively detects skewed or missing values before training begins?

A.Manually inspecting the first 100 rows of the dataset using a notebook cell.
B.Using DLT Expectations to define and enforce constraints on input data.
C.Relying on the model to handle missing values by using imputation during training.
D.Increasing the size of the training cluster to process more data.
AnswerB

DLT Expectations provide a declarative way to define data quality rules. By automatically monitoring data against these rules and failing the pipeline or isolating bad records, engineers ensure that only high-quality data enters the training set, which is essential for building reliable, production-grade machine learning systems.

Why this answer

Integrating data quality checks, such as those provided by Delta Live Tables (DLT) expectations or Great Expectations, is crucial for MLOps. By enforcing schema and statistical constraints at the ingest and preparation stages, the pipeline can fail early if data quality falls below standards, preventing the training of 'garbage-in-garbage-out' models and saving compute costs while maintaining model reliability in production.

Exam trap

Candidates often suggest post-training model monitoring tools for pre-training data issues, failing to recognize that DLT Expectations are designed specifically for proactive data validation during the ETL pipeline phase.

142
MCQmedium

Refer to the exhibit. A data scientist is logging a Scikit-Learn model to the MLflow Model Registry. Which benefit does providing the `signature` and `input_example` offer during the deployment phase?

A.It automatically scales the cluster size based on the input example size.
B.It enables MLflow to enforce data types during inference, preventing schema mismatch errors.
C.It allows the model to be trained in distributed mode using PySpark.
D.It automatically converts the model into an ONNX format for faster inference.
AnswerB

The model signature acts as a contract between the model and the caller. If the provided input does not match the signature's types, the model serving endpoint will reject the request with a clear error, ensuring that the inference engine receives the exact data format the model expects.

Why this answer

Providing a model signature and input example defines the expected data schema, allowing MLflow to perform type validation during inference. This is vital for production systems, as it prevents runtime errors caused by malformed inputs and enables Databricks to automatically generate deployment documentation and test payloads, significantly reducing the debugging time when deploying models to Model Serving endpoints.

Exam trap

Candidates frequently confuse model signatures with performance evaluation metrics, incorrectly thinking signatures measure accuracy rather than enforcing strict input and output data types.

143
MCQhard

A company has deployed a model to a Databricks Model Serving endpoint. The model's predictions must be logged to a Delta table for monitoring and auditing. The ML engineer wants to enable inference logging without modifying the model's code. Which approach achieves this with minimal effort?

A.Use a Databricks job to periodically query the endpoint's logs via the REST API and insert them into a Delta table.
B.Wrap the model's predict method to write inputs and outputs to a Delta table using Spark, then redeploy the model.
C.Configure the model to log its predictions to MLflow, then enable MLflow tracking for the serving endpoint.
D.Enable inference logging on the serving endpoint by specifying a Delta table path in the endpoint configuration.
AnswerD

Databricks Model Serving supports inference logging, which automatically captures request and response payloads and writes them to a specified Delta table. This can be enabled in the endpoint configuration without changing the model code. It is the intended feature for auditing and monitoring, and requires only setting the logging destination.

Why this answer

Inference logging is a built-in feature of Databricks Model Serving that automatically logs request and response data to a Delta table. It can be enabled via the endpoint configuration without altering the model code. Other options require code changes, custom jobs, or misuse of MLflow tracking, and do not provide the same seamless auditing capability.

Exam trap

The trap here is assuming that MLflow tracking can be used for inference logging, or that manual code changes are necessary, when Databricks provides a native inference logging feature.

144
MCQeasy

A data scientist has deployed a model to a Databricks Model Serving endpoint. The endpoint is configured with scale-to-zero enabled and a workload size of Small. After a period of inactivity, the endpoint scales down to zero. A client application sends a request to the endpoint after this idle period. What happens to the first request?

A.The request fails with a 503 Service Unavailable error because the endpoint is offline.
B.The request is immediately served by a cold-start replica with no added latency.
C.The request is queued until the endpoint scales up, then processed with increased latency.
D.The request is routed to a fallback model version that is always kept warm.
AnswerC

When scale-to-zero is enabled, the endpoint scales down to zero replicas after inactivity. The first request after idle time triggers a scale-up, causing the request to be queued until a replica is available. This results in higher latency for that initial request. Subsequent requests are served with normal latency.

Why this answer

With scale-to-zero, the endpoint reduces replicas to zero after inactivity. When a new request arrives, the system must provision resources and load the model, causing the request to be queued and served with higher latency. This is expected behavior and not an error.

Exam trap

The trap here is assuming that scale-to-zero causes request failures or that a warm fallback exists, when in fact the request is queued and served after a cold start.

145
MCQeasy

An ML engineer has deployed a model to Databricks Model Serving and wants to monitor the endpoint's performance over time. They need to track the number of requests, latency, and error rates. Which Databricks feature provides these metrics out-of-the-box?

A.Delta Live Tables
B.Databricks Model Serving endpoint metrics in the Databricks UI
C.Unity Catalog audit logs
D.MLflow Tracking
AnswerB

Databricks Model Serving provides built-in metrics such as request count, latency, and error rates, which are displayed in the endpoint's detail page in the Databricks UI. These metrics are available without additional configuration and can be used to monitor the health and performance of the endpoint.

Why this answer

Databricks Model Serving includes built-in metrics that are accessible in the Databricks UI. These metrics cover request counts, latency, and error rates, providing immediate visibility into endpoint performance. Other options like MLflow Tracking or Unity Catalog audit logs serve different purposes and do not offer out-of-the-box endpoint monitoring.

Exam trap

The trap here is assuming MLflow Tracking, which is used during training, also monitors deployed endpoints, when in fact Model Serving has its own metrics dashboard.

146
MCQhard

An ML engineer is updating a production model serving endpoint to use a new model version. The endpoint currently serves version 1 with the 'Champion' alias. The engineer wants to test version 2 with a small percentage of live traffic before full rollout. Which deployment strategy should they use in Databricks Model Serving?

A.Create a new endpoint for version 2 and use a load balancer to split traffic.
B.Configure traffic splitting on the existing endpoint to route a percentage to version 2.
C.Update the 'Champion' alias to point to version 2 and monitor performance.
D.Deploy version 2 to a staging endpoint and compare metrics offline.
AnswerB

Databricks Model Serving supports serving multiple model versions within a single endpoint by specifying a traffic split percentage for each version. This allows canary deployments where a small portion of traffic goes to the new version while the majority remains on the stable version. It is the recommended approach for testing new versions with live traffic.

Why this answer

Databricks Model Serving allows you to serve multiple model versions on a single endpoint and split traffic between them by specifying percentages. This enables a canary deployment where a small fraction of requests go to the new version, allowing real-world testing before a full rollout. Other options either cause a full cutover, require external tools, or do not use live traffic.

Exam trap

The trap here is thinking that updating the alias or creating a separate endpoint is the way to do canary testing, but Databricks provides native traffic splitting within an endpoint.

147
MCQeasy

A data scientist registers a new model version to Unity Catalog and wants to promote it to Production after validation. Which action accomplishes this in the current Databricks recommendation?

A.Set the model version's stage to Production using the MLflow Model Registry stage transition API.
B.Assign the alias 'champion' to the validated model version using the MLflow client or Catalog Explorer.
C.Add a tag named 'stage' with the value 'Production' to the model version and restart the serving endpoint.
D.Create a new registered model with the same name plus a '-prod' suffix and copy the artifacts into it.
AnswerB

Databricks recommends using aliases such as 'champion' or 'Production' to mark the version that serving endpoints and jobs should consume. Assigning the alias to the validated version promotes it without mutating the version itself, and endpoints configured to serve that alias pick up the change.

Why this answer

In Unity Catalog, model promotion is done with aliases rather than the deprecated stages. Assigning an alias such as 'champion' to a validated version marks it as the one that serving endpoints and jobs should consume, and endpoints referencing that alias automatically serve the newly promoted version without changing their configuration.

Exam trap

The trap here is assuming that MLflow stages still govern promotion in Unity Catalog, when aliases are the supported mechanism.

148
MCQhard

A financial institution has deployed a credit risk model to a Databricks Model Serving endpoint. The model was trained on data that includes sensitive customer attributes. The compliance team requires that all predictions be explainable and that the model's decisions can be audited. The data science team wants to use SHAP (SHapley Additive exPlanations) to generate explanations for each prediction. Which approach should they take to integrate SHAP with the serving endpoint while maintaining low latency?

A.Enable automatic logging of SHAP values by setting the `log_explainer` parameter in the MLflow model signature.
B.Use the `shap.Explainer` within a custom PyFunc model that computes SHAP values on the fly and returns them alongside predictions.
C.Deploy a separate endpoint that runs SHAP explanations asynchronously and return a job ID that the client can poll for results.
D.Precompute SHAP values for all possible input combinations and store them in a Delta table for lookup during serving.
AnswerB

Wrapping the model in a custom PyFunc that computes SHAP values at inference time allows explanations to be generated for each request. While this adds latency, it can be optimized by using efficient SHAP implementations and limiting the number of background samples. This approach provides real-time explanations and can be deployed to Model Serving. It meets the requirement for per-prediction explainability and auditability, though latency must be managed.

Why this answer

Integrating SHAP into a custom PyFunc model allows the serving endpoint to return both predictions and explanations in real time. This satisfies the compliance need for per-prediction explainability and auditability. While it introduces some latency, it can be optimized with efficient SHAP algorithms and careful resource allocation, making it suitable for real-time serving in regulated environments.

Exam trap

The trap here is assuming that SHAP explanations can be precomputed or automatically logged, when in reality they must be computed at inference time for each request to provide accurate, real-time explanations.

149
MCQeasy

Which of the following describes the 'Gold' layer in the Medallion Architecture, and why is it important for machine learning?

A.It contains raw data ingested directly from external sources.
B.It contains aggregated, business-level data ready for model training.
C.It is used to store model artifacts and logs.
D.It is where data scientists perform exploratory data analysis.
AnswerB

Gold tables provide clean, validated data that matches business requirements. This makes them the ideal source for training models. By building models on Gold data, scientists ensure that their features are based on reliable information, significantly reducing the probability of errors caused by poor data quality in production pipelines.

Why this answer

The Gold layer contains refined, business-level data that is ready for consumption. In MLOps, it is the standard source for training data. By using curated, high-quality data from the Gold layer, data scientists avoid the noise and inconsistencies of raw data, leading to more robust models and faster iteration times, as they spend less time on manual data cleaning and validation.

Exam trap

Candidates often confuse 'Gold' data with 'Silver' data, failing to realize that Gold specifically implies business-level aggregation, which is the necessary state for final model training inputs.

150
MCQhard

A machine learning engineer is using MLflow to track experiments and wants to compare multiple runs to identify the best model. They have logged metrics such as accuracy, precision, and recall. Which MLflow feature allows them to programmatically retrieve and compare these metrics across runs for further analysis?

A.mlflow.log_metric()
B.mlflow.search_runs()
C.mlflow.list_experiments()
D.mlflow.get_run()
AnswerB

mlflow.search_runs() allows you to query runs across experiments using a SQL-like filter and returns a pandas DataFrame containing metrics, parameters, and tags for all matching runs. This makes it easy to programmatically compare metrics across runs, sort them, and perform further analysis. It is the most efficient way to retrieve and compare multiple runs in MLflow.

Why this answer

mlflow.search_runs() is designed to query and retrieve run data across experiments. It returns a pandas DataFrame with metrics, parameters, and tags, enabling programmatic comparison and analysis. This is ideal for identifying the best model by sorting and filtering based on metrics like accuracy or precision.

Exam trap

The trap here is confusing functions that retrieve a single run or list experiments with the one that retrieves multiple runs for comparison.

Page 1

Page 2 of 4

Page 3

All pages

Practice Databricks-ML-Pro by domain

Target a specific domain to shore up weak areas.

See all domains with question counts →