Courseiva

Databricks Certified Machine Learning Professional (Databricks-ML-Pro) — Questions 226–300

300 questions total · 4pages · All types, answers revealed

Page 3

Page 4 of 4

226
MCQmedium

A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. The feature table contains a column `transaction_time` that is a timestamp. After creating the training set with `create_training_set`, the resulting DataFrame includes `transaction_time` but the model training code fails because the timestamp is not accepted by the XGBoost trainer. What is the most likely cause and correct resolution?

A.The training set includes the timestamp column because `create_training_set` does not automatically exclude non-numeric columns; the data scientist should drop or transform the timestamp column before passing the DataFrame to the XGBoost trainer.
B.The timestamp column should be excluded from the feature table and instead passed as a separate DataFrame column outside the Feature Store lookup.
C.The error occurs because the Feature Store requires all features to be of type double; the data scientist must cast the timestamp to double using `.cast('double')` before creating the training set.
D.Databricks Feature Store automatically converts timestamp columns to Unix epoch integers when creating the training set, so the trainer should accept them without modification.
AnswerA

`create_training_set` returns all columns from the feature table, including timestamps, without altering their types. XGBoost requires numeric or boolean inputs, so a raw timestamp causes a failure. The correct fix is to drop the timestamp or derive numeric features such as hour of day or day of week. This preserves feature lineage while making the data compatible with the trainer.

Why this answer

`create_training_set` preserves the original data types of feature columns, so a timestamp column remains a timestamp. XGBoost cannot handle timestamp types directly, causing the training failure. The data scientist must either drop the column or transform it into numeric features such as hour, day of week, or time since a reference point.

This ensures compatibility while retaining useful temporal signals.

Exam trap

The trap here is assuming that Databricks Feature Store automatically converts non-numeric columns like timestamps into numeric formats suitable for all model trainers.

227
MCQmedium

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires a custom pre-processing step that is not part of the standard MLflow transformers or pyfunc flavor. Which deployment approach ensures the custom logic executes reliably within the serverless serving container?

A.Register the vanilla PyTorch state dict and apply pre-processing transformations inside the client application before sending HTTP payloads.
B.Define a separate Spark UDF inside the serving endpoint configuration file to intercept incoming JSON batches.
C.Implement a custom MLflow pyfunc PythonModel subclass that encapsulates both the PyTorch model and the custom pre-processing transformations, then log it with MLflow.
D.Store the pre-processing code in a separate volume and reference its absolute file path in the model serving endpoint environment variables.
AnswerC

Subclassing MLflow PythonModel allows developers to bundle custom inference logic, tokenizers, or scalers directly into the logged artifact. Databricks Model Serving natively understands the pyfunc flavor, executing the overridden predict method securely within the managed container environment.

Why this answer

Packaging the custom pre-processing logic directly into the MLflow pyfunc model artifact by overriding the predict context ensures that all required transformations travel with the model weights. This guarantees identical execution behavior between local testing and production serverless endpoints without relying on external pipeline code.

Exam trap

Candidates rely on default MLflow model flavors for custom architectures, forgetting that non-standard pre-processing steps require custom wrapper logic to execute inside serverless endpoints.

228
MCQmedium

A data scientist is using the Databricks Feature Store to build a training set for a fraud-detection model. The feature table is created with a primary key of `customer_id` and a timestamp key of `transaction_ts`. When calling `create_training_set`, the scientist wants to ensure that each label row receives exactly the most recent feature value available at or before the label's timestamp. Which argument must be supplied to `create_training_set` to enforce this point-in-time behavior?

A.Set `timestamp_lookup_key` to the label DataFrame's event timestamp column.
B.Set `label` to the timestamp column of the label DataFrame.
C.Set `exclude_columns` to the timestamp key so the join ignores time.
D.Set `feature_names` to the timestamp key of the feature table.
AnswerA

Supplying `timestamp_lookup_key` tells the Feature Store which column in the label DataFrame represents the event time, so it performs a point-in-time lookup and joins only feature values whose timestamp key is less than or equal to that value. This guarantees no future leakage into the training set, which is essential for a realistic fraud-detection model.

Why this answer

Point-in-time correctness in the Databricks Feature Store is achieved by telling the training-set builder which label column holds the event timestamp. That column is passed as `timestamp_lookup_key`, and the Feature Store then joins only feature rows whose timestamp key precedes or equals it. Without this argument, the join can pull the latest feature value regardless of time, introducing label leakage that inflates offline metrics and misleads deployment decisions.

Exam trap

The trap here is assuming any timestamp-related argument enables point-in-time joins, when only the dedicated timestamp lookup key parameter controls that behavior.

229
MCQmedium

You are performing hyperparameter tuning using Hyperopt on Databricks. Which TWO configurations must be defined to ensure optimal performance and result tracking?

A.Use SparkTrials to distribute the hyperparameter search across the cluster.
B.Call mlflow.end_run() manually at the end of every trial function.
C.Wrap the objective function in an MLflow autologging block.
D.Log hyperparameters and metrics explicitly inside the objective function.
E.Set the search space to use a linear scale for all numerical hyperparameters.
AnswerA, D

SparkTrials is essential for parallelizing hyperparameter tuning jobs on Databricks. It manages the distribution of trial tasks across available worker nodes, preventing bottlenecks that occur when executing sequential trials on a single driver node, which is inefficient for large-scale training pipelines.

Why this answer

Using SparkTrials allows Hyperopt to distribute trial execution across multiple cluster nodes, significantly accelerating grid or random search. Integrating MLflow within the objective function ensures that each trial's parameters, metrics, and artifacts are captured, enabling users to analyze the tuning history and select the best model based on validated performance metrics.

Exam trap

Candidates often select standard Hyperopt without SparkTrials, assuming parallelization happens automatically, or forget that metrics must be explicitly logged inside the objective function to be tracked properly.

230
MCQmedium

An ML engineer needs to deploy a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default Databricks Runtime and must be installed from a private PyPI repository. Which approach should be used to include this package in the model's environment?

A.Install the package on the driver node of the serving cluster using a %pip magic command before deployment.
B.Add the package to the cluster's init script and reference the cluster when creating the serving endpoint.
C.Upload the package as a Databricks notebook and import it at runtime within the model's predict function.
D.Include the package in the model's conda.yaml or requirements.txt file when logging the model with MLflow.
AnswerD

When logging a model with MLflow, you can specify dependencies in a conda.yaml or requirements.txt file. Databricks Model Serving uses these files to build the serving container, installing the listed packages. By including the private PyPI package with its index URL in the requirements, the serving environment will have the necessary dependency, ensuring the model runs correctly.

Why this answer

Databricks Model Serving builds a container for the model based on the environment specified during MLflow model logging. By including the private PyPI package in the conda.yaml or requirements.txt with the appropriate index URL, the package is installed in the serving environment. This is the supported method for custom dependencies.

Exam trap

The trap here is assuming that serving endpoints can use cluster init scripts or interactive pip commands, which are not available in the managed serving environment.

231
MCQmedium

Refer to the exhibit. A data scientist is preparing to log a model. What is the primary benefit of including the explicit 'signature' provided in the exhibit during the mlflow.log_model process?

A.It enables automatic hyperparameter tuning during the next training iteration.
B.It allows the serving endpoint to perform automatic data type validation on incoming requests.
C.It automatically encrypts the model weights for secure storage in the registry.
D.It forces the model to be converted into a serialized format like ONNX.
AnswerB

Defining a signature allows Databricks Model Serving to validate that the schema of incoming request payloads matches the expected input structure. If a request contains incorrect data types, the service rejects it immediately, preventing downstream processing failures and providing clear error messages for debugging integration issues.

Why this answer

The model signature acts as a contract between the model and its environment. By explicitly defining the input and output schema, Databricks can perform validation during inference to prevent runtime errors caused by mismatched data types. This is critical for production pipelines where data quality might fluctuate, ensuring that the model consumes and produces the expected data formats reliably.

Exam trap

Candidates assume the signature is purely for documentation or metadata purposes. They fail to realize it is a functional requirement for the serving endpoint to perform automatic runtime data validation.

232
MCQmedium

A data science team is using Databricks Model Serving to deploy a model that must process sensitive data. They need to ensure that all inference requests are logged for auditing purposes, including the input data and predictions. Which approach should they take?

A.Use MLflow Tracking to log each inference request as a run in an experiment.
B.Set up a Databricks Job that periodically queries the endpoint and logs the responses.
C.Enable inference logging on the Model Serving endpoint and configure a Delta table to store the logs.
D.Configure the model to write logs to a file in DBFS using a custom Python logger.
AnswerC

Databricks Model Serving supports inference logging, which captures request and response payloads to a Delta table. This feature is designed for auditing and monitoring, allowing you to store input data and predictions. By enabling it and specifying a Delta table, the team can meet compliance requirements and analyze model performance over time.

Why this answer

Inference logging in Databricks Model Serving is the built-in feature for capturing request and response data to a Delta table. It is designed for auditing, monitoring, and compliance, providing a scalable and managed solution. Other methods like custom logging or MLflow Tracking are not suited for production-scale inference logging and may introduce reliability or performance issues.

Exam trap

The trap here is assuming that any logging mechanism (like MLflow or custom code) can serve the purpose, without recognizing that Model Serving has a dedicated inference logging feature that simplifies compliance.

233
MCQmedium

What is the primary advantage of using a Model-as-Code approach in Databricks for machine learning deployments?

A.It ensures that the model is always trained on the latest data.
B.It enables consistent, reproducible, and versioned deployment environments.
C.It automatically generates unit tests for all machine learning code.
D.It eliminates the need for any monitoring of model performance.
AnswerB

By codifying everything from environment settings to deployment logic, teams ensure that the production environment is identical to the testing environment. This consistency is the foundation of reliable MLOps, enabling teams to deploy with confidence and revert to previous known-good states if issues arise in production.

Why this answer

Model-as-Code treats the entire deployment process—infrastructure, environment configuration, and code—as versioned artifacts. This allows for total reproducibility, where any production state can be rolled back or recreated using stored definitions. This practice is essential for enterprise MLOps, as it ensures that deployments are predictable, scalable, and audit-compliant, significantly reducing the risks associated with manual configuration changes in a production environment.

Exam trap

Candidates often describe Model-as-Code as merely 'automating deployments,' failing to emphasize the critical aspects of reproducibility, environment parity, and versioning that define the approach.

234
MCQmedium

You are developing a machine learning pipeline where you need to perform feature engineering on a large dataset using Spark, then train a model using Scikit-Learn. Which workflow is most efficient?

A.Convert the raw Spark DataFrame to Pandas immediately, then perform feature engineering.
B.Perform feature engineering using Spark transformations, then collect to Pandas for training.
C.Use a custom UDF to execute Scikit-Learn code on every partition of the Spark DataFrame.
D.Rewrite all feature engineering logic using only Scikit-Learn transformers.
AnswerB

This method optimizes resource usage by using Spark for parallel processing of large-scale data. Once the data is transformed and sufficiently small, collecting it to the driver memory enables the use of Scikit-Learn, which is optimized for in-memory single-node training, balancing performance with library capabilities.

Why this answer

The most efficient workflow involves leveraging Spark for distributed data transformation and feature engineering, followed by collecting the processed data into a Pandas DataFrame for Scikit-Learn training. This approach uses the strengths of both frameworks, ensuring that computationally expensive transformations are parallelized across the cluster, while Scikit-Learn handles the specific modeling logic that is not natively distributed.

Exam trap

Candidates often try to perform all transformations inside Scikit-Learn or attempt to run Spark on the entire Scikit-Learn training process, failing to split the workflow based on framework strengths.

235
MCQmedium

A machine learning team is using `mlflow.autolog()` to track experiments. They notice that certain custom metrics are not being captured. What is the most effective way to address this?

A.Disable autologging and manually log every single parameter and metric.
B.Implement a custom callback in the training loop that uses mlflow.log_metric for the specific metrics.
C.Increase the logging frequency of the autologger via the configuration file.
D.Create a custom Python class that inherits from the MLflow Autologger base class.
AnswerB

Using standard MLflow logging functions within a training loop allows for the integration of domain-specific metrics alongside the automatically logged framework metrics. This approach provides maximum flexibility, ensuring that all necessary data for model evaluation is captured in a single, unified experiment run record for analysis.

Why this answer

While autologging captures standard metrics like accuracy, custom metrics require explicit logging calls using `mlflow.log_metric`. This is important because business-specific KPIs, such as profit margin or specific domain-based error rates, often require custom logic that framework-specific autologgers cannot infer automatically. Explicit logging ensures comprehensive experiment visibility and allows for precise model selection based on business goals rather than just technical performance.

Exam trap

Candidates mistakenly believe that enabling `mlflow.autolog()` is sufficient to capture all domain-specific or custom business KPIs without writing explicit logging statements.

236
MCQhard

A machine learning team is using MLflow Model Registry to manage a model that is deployed to a production endpoint. They need to implement a CI/CD pipeline that automatically transitions a model version from 'Staging' to 'Production' only after it passes a set of validation tests. Which MLflow feature allows them to trigger the transition based on test results?

A.MLflow Model Registry stage transitions via the REST API
B.MLflow Projects with conditional steps
C.Databricks Jobs with task dependencies
D.MLflow Model Registry webhooks
AnswerA

The MLflow Model Registry REST API provides endpoints to transition model versions between stages. A CI/CD pipeline can call these endpoints programmatically after validation tests pass. This allows automated promotion based on test outcomes. The API is the standard way to integrate stage transitions into external workflows.

Why this answer

The MLflow Model Registry REST API allows programmatic stage transitions. In a CI/CD pipeline, after validation tests pass, you can call the API to move the model version from Staging to Production. This integrates seamlessly with automation tools.

Webhooks are for notifications, not for initiating transitions, and other options lack the direct capability to change stages based on test results.

Exam trap

The trap here is confusing webhooks, which react to transitions, with the API that actually performs the transition on demand.

237
MCQmedium

You are responsible for a fraud detection model deployed to a Databricks Model Serving endpoint. The model was trained on transaction data from the past 6 months. After two months in production, you notice a gradual decline in precision and recall. You suspect data drift. Which approach should you use to monitor feature drift for this endpoint?

A.Set up a Databricks SQL dashboard that queries the endpoint's logs and calculates the average prediction value over time.
B.Configure the endpoint to automatically retrain the model when drift is detected using the built-in auto-retrain feature.
C.Use MLflow Tracking to log the model's predictions and manually compare them to ground truth labels in a notebook.
D.Enable inference logging on the endpoint and use Databricks Lakehouse Monitoring to compare production feature distributions against the training baseline.
AnswerD

Inference logging captures the feature values sent to the endpoint, and Lakehouse Monitoring can compute drift metrics by comparing these to a baseline profile from training data. This provides automated drift detection and alerts, directly addressing the scenario.

Why this answer

Inference logging captures the actual feature values sent to the model endpoint, and Lakehouse Monitoring provides built-in drift metrics by comparing these to a baseline profile. This combination enables automated, scalable drift detection without manual intervention. Other options either misrepresent platform features or fail to statistically compare distributions.

Exam trap

The trap here is assuming that Model Serving includes automatic retraining or that simple prediction monitoring suffices for drift detection.

238
Multi-Selectmedium

A data science team is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. They want to configure auto-scaling appropriately. Which TWO parameters should they adjust to control the scaling behavior? (Choose two.)

Select 2 answers
A.Enable scale-to-zero to reduce costs during idle periods.
B.Adjust the model's batch size in the predict function to process more requests per inference.
C.Set the minimum number of replicas to a value greater than zero.
D.Configure the maximum number of replicas to a high value to allow scaling out.
E.Set the endpoint's workload size to 'Large' to increase per-replica capacity.
AnswersC, D

Setting a minimum replica count ensures that a baseline capacity is always available, reducing cold start impact and allowing the endpoint to handle initial bursts without waiting for scaling. This is crucial for latency-sensitive applications that cannot tolerate cold starts.

Why this answer

Auto-scaling in Databricks Model Serving is governed by the minimum and maximum replica counts. The minimum ensures baseline capacity to handle initial bursts, while the maximum allows the endpoint to scale out to meet high demand. These two parameters directly control the scaling range and are essential for handling traffic spikes without dropping requests.

Exam trap

The trap here is focusing on cost-saving features like scale-to-zero, which actually hinder burst handling, instead of the core scaling parameters.

239
MCQmedium

You are implementing model monitoring to detect data drift. Which Databricks feature should you use to automatically track and alert on changes in the distribution of input data features over time?

A.MLflow Experiment Tracking
B.Delta Live Tables Expectations
C.Databricks Lakehouse Monitoring
D.Databricks SQL Alerts
AnswerC

Lakehouse Monitoring is the purpose-built service for detecting drift in input features and model predictions. It automatically creates dashboards and alerts by comparing production data against training baselines, providing essential visibility into model health after deployment in a production environment.

Why this answer

Databricks Lakehouse Monitoring provides automated insights into data quality and drift. By tracking statistical properties of incoming data against a baseline, it generates alerts when significant shifts are detected. This is a crucial component of MLOps, as it proactively identifies when a model's performance might degrade due to environmental changes, allowing teams to retrain or adjust models before they negatively impact business outcomes.

240
MCQhard

Refer to the exhibit. A machine learning team has updated their model serving endpoint configuration as shown in the JSON. Which deployment strategy is being implemented, and what is the primary risk associated with this specific configuration?

A.Blue/Green deployment; the primary risk is the high cost of running two identical clusters simultaneously.
B.Canary deployment; the primary risk is increased 'cold start' latency for the v2 model due to low traffic volume.
C.A/B testing; the primary risk is that the workload size 'Small' is insufficient for 90% of the traffic.
D.Shadow deployment; the primary risk is that v2 will interfere with the predictions returned by v1.
AnswerB

In a Canary setup with a 90/10 split and scale-to-zero enabled, the v2 instance will likely idle frequently. When the 10% of requests do arrive, the system must provision the instance from scratch, causing significant delays for those specific users compared to the more frequently used v1 version.

Why this answer

The exhibit demonstrates a Canary deployment where a small fraction of traffic (10%) is routed to a new model version (v2) while the majority remains on the stable version (v1). This allows for real-world testing with minimal impact. However, since both versions are set to scale-to-zero, the 10% traffic might not be frequent enough to keep v2 warm, leading to high latency for those users.

Exam trap

Candidates often overlook the 'cold start' penalty in Canary deployments. They assume that if it works for 10% of traffic, it works for everything, forgetting the resource initialization latency.

241
MCQmedium

Which practice is most effective for managing dependencies to ensure consistent model training and inference results across different Databricks clusters?

A.Manually install libraries on each cluster node before starting the job.
B.Include a requirements.txt file with the notebook or project and specify it in the MLflow model logging process.
C.Rely on the default Databricks Runtime library set for all projects.
D.Use broad version constraints like '>=1.0' in the environment configuration.
AnswerB

Including a requirements file ensures that the model environment is explicitly captured as part of the model artifact. When deployed, the serving platform can recreate the exact environment, ensuring that the inference code runs against the same dependency versions used during training, which minimizes numerical inconsistencies and runtime errors.

Why this answer

Using environment management files (e.g., requirements.txt or conda.yaml) ensures that the exact library versions used during development are replicated in the training and serving environments. This consistency is critical for preventing 'works on my machine' issues, where discrepancies in library versions lead to different numerical outputs or runtime crashes during inference, thus maintaining the reliability and reproducibility of the machine learning pipeline.

Exam trap

Candidates often assume manual library installation via %pip install in the notebook is sufficient, ignoring that these changes do not persist across different clusters or automated production deployment jobs.

242
MCQmedium

A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. They define a feature table with a primary key of `transaction_id` and a timestamp key of `event_ts`. When creating the training set with `create_training_set`, they specify `lookup_key=['transaction_id']`. The resulting training set contains features from multiple feature tables. Which statement describes how point-in-time correctness is ensured during this operation?

A.The `create_training_set` function automatically sorts features by their creation date and selects the latest values before the label timestamp.
B.The timestamp key is automatically used to join features as of the timestamp of each label event, preventing data leakage from future feature values.
C.The primary key alone guarantees point-in-time correctness because it uniquely identifies each transaction and its associated features.
D.Point-in-time correctness is enforced only if the feature tables are registered with a `timestamp` column that matches the label DataFrame's index.
AnswerB

Databricks Feature Store uses the timestamp key to perform time-series joins, ensuring that for each label event, only feature values with timestamps at or before the label timestamp are included. This prevents leakage of future information and maintains point-in-time correctness, which is critical for fraud detection where future transactions could otherwise contaminate the training data.

Why this answer

Point-in-time correctness in Databricks Feature Store is achieved by time-series joins using the timestamp key. When creating a training set, the system matches each label row with feature values whose timestamps are less than or equal to the label timestamp. This ensures that only historical, non-leaking features are used.

The primary key alone is insufficient; the timestamp key is essential for temporal alignment.

Exam trap

The trap here is assuming that a primary key join automatically prevents data leakage, when in fact the timestamp key is required to enforce point-in-time correctness.

243
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that uses a custom transformer requiring a GPU. The endpoint must handle high throughput with low latency. Which workload type and configuration should be selected?

A.Use a 'Large' workload type and specify GPU requirements in the model's conda environment.
B.Use a 'Medium' workload type and configure the endpoint to use GPU by setting an environment variable.
C.Use a 'Small' workload type with CPU and enable GPU acceleration via a model parameter.
D.Use a 'GPU_SMALL' or 'GPU_MEDIUM' workload type, depending on the model's memory and compute needs.
AnswerD

Databricks Model Serving offers GPU-enabled workload types such as 'GPU_SMALL' and 'GPU_MEDIUM'. These provide GPU instances suitable for models that require GPU acceleration. Selecting the appropriate size based on memory and compute requirements ensures high throughput and low latency for GPU-dependent models.

Why this answer

Databricks Model Serving provides specific GPU-enabled workload types, such as 'GPU_SMALL' and 'GPU_MEDIUM', which include GPU instances. For a model requiring GPU, selecting one of these workload types is necessary. The choice between small and medium depends on the model's memory and compute demands to achieve high throughput and low latency.

Exam trap

The trap here is assuming that GPU can be enabled through environment variables or conda specifications, when in fact it must be selected as a workload type during endpoint configuration.

244
MCQmedium

Your organization requires that all models deployed to production must be signed by a security officer. How can you enforce this requirement within the Databricks MLflow Model Registry?

A.Delete all models in the Production stage and only re-upload the signed models.
B.Grant 'CAN_MANAGE' permissions on the registry to everyone to ensure transparency.
C.Restrict 'CAN_MANAGE' permissions on the model, and use an automated CI/CD pipeline for promotion.
D.Set the model version description to 'Signed by Security' after manual inspection.
AnswerC

Restricting permissions ensures that only authorized service principals or security officers can manage transitions. By requiring an automated CI/CD pipeline, you ensure that the promotion process is audited and validated against security policies before the model is moved to the Production stage, effectively enforcing the organization's requirements.

Why this answer

Enforcing approval workflows is crucial for compliance and risk management in MLOps. By using MLflow's permissions and stages, you can restrict who can promote models to 'Production'. Combining this with programmatic checks or external CI/CD gates ensures that a model cannot be deployed without the necessary authorization, preventing unauthorized or unvetted code from reaching production inference services.

Exam trap

Candidates often suggest using manual UI approvals, ignoring the requirement for automated, audit-compliant CI/CD pipelines that enforce security officer sign-offs through programmatic controls.

245
MCQmedium

A team uses MLflow Projects to package training code and runs jobs on Databricks clusters. They want to ensure that a job run today can be reproduced six months later with the same library versions, even if the cluster's base image and PyPI packages have changed. Which practice best achieves this?

A.Pin exact library versions in the project's conda.yaml or requirements.txt, and log the resulting environment specification with the run.
B.Rely on the Databricks Runtime version of the cluster, since Databricks Runtime images are immutable and never receive library updates.
C.Store the trained model artifacts in DBFS and re-run the training notebook from the workspace revision history.
D.Use the latest versions of all libraries at run time and rely on MLflow autologging to record which versions were used.
AnswerA

Pinning exact versions in the project environment file ensures the dependency resolver installs the same libraries at run time, and logging the environment specification with the run captures what was actually used. This combination makes the run reproducible months later regardless of changes to the base image or PyPI, because the pipeline can rebuild the same environment from the recorded specification rather than relying on floating versions.

Why this answer

Reproducibility across time requires controlling the environment, not just observing it. Pinning exact versions in the project's environment file guarantees the same dependencies are installed, and logging the environment specification with the run records the actual resolved environment for future rebuilding. Relying on runtime versions, notebook revisions, or autologging alone does not prevent dependency drift, so those approaches cannot ensure the same result six months later.

Exam trap

The trap here is confusing recording the environment with controlling it, since autologging captures versions but does not pin them for future runs.

246
MCQmedium

A machine learning engineer is using Databricks Model Serving to deploy a model that requires a custom Python library not available in the default environment. The engineer wants to ensure the endpoint uses the exact library version and that the deployment is reproducible. Which approach should the engineer take?

A.Install the custom library on all driver and worker nodes of the cluster used for model training, then deploy the model to Model Serving.
B.Use the Databricks REST API to upload the library to the Model Serving endpoint after deployment, then restart the endpoint.
C.Package the custom library as a Python wheel and include it in the model's conda environment or requirements file when logging the model with MLflow.
D.Create an init script that installs the custom library on the Model Serving cluster, and attach it to the endpoint configuration.
AnswerC

MLflow models can capture dependencies via a conda environment or requirements file. By packaging the custom library as a wheel and including it in the model's environment specification, Databricks Model Serving will install that exact version when deploying the model. This ensures reproducibility and availability of the custom library at inference time.

Why this answer

To ensure a custom library is available and reproducible in Databricks Model Serving, it must be included as a dependency when logging the MLflow model. Packaging the library as a wheel and adding it to the conda environment or requirements file ensures that Model Serving installs the exact version during deployment. This is the supported and recommended practice.

Exam trap

The trap here is assuming that libraries installed on a training cluster or via init scripts will carry over to the serverless Model Serving environment, which is not the case.

247
MCQmedium

Refer to the exhibit. Your automated CI/CD pipeline triggered a model deployment to production, but the job failed with the error shown. What is the most likely cause?

A.The model was registered, but the cluster lacks permissions to read it.
B.The deployment job is referencing a model version that hasn't been created yet.
C.The model serving endpoint is already occupied by another model version.
D.The training cluster ran out of memory during the model artifact upload.
AnswerB

The REST exception confirms the system attempted to fetch a non-existent entity. In CI/CD, this suggests a dependency failure where the deployment script executes before the training job successfully completes the registration, indicating a need for improved job orchestration or explicit completion gating.

Why this answer

The error indicates that the deployment pipeline is attempting to access a model version that does not exist in the Model Registry. This usually happens when the pipeline triggers before the registration process completes or when a race condition occurs between the training job and the deployment job. Validating the existence of the model version before initiating deployment is essential for pipeline reliability.

Exam trap

Candidates often blame infrastructure or network issues. However, in automated CI/CD, the most common failure is a race condition where the deployment script executes before the registration job finishes.

248
MCQeasy

What is the primary purpose of registering a model in the MLflow Model Registry?

A.To increase the training speed of the machine learning model.
B.To provide a structured workflow for versioning, stage management, and model governance.
C.To automatically retrain the model when data drift is detected.
D.To convert Python code into high-performance C++ code.
AnswerB

The Registry allows teams to manage the lifecycle of models by tracking versions and facilitating transitions between stages like 'Staging' and 'Production'. This ensures that only validated models are deployed, providing an audit trail and reducing the risk of unauthorized or unverified changes reaching the production environment.

Why this answer

The MLflow Model Registry provides a centralized hub for managing the model lifecycle, including versioning, stage transitions (e.g., Staging to Production), and lineage tracking. This is foundational for MLOps because it provides a single source of truth for all stakeholders, enabling controlled deployments, auditability, and the ability to easily revert to previous model versions if performance regressions are detected in production.

Exam trap

Candidates confuse the Model Registry with the MLflow Tracking server, mistakenly believing the registry is for storing raw experiment metrics rather than managing deployment stages and model lifecycle versions.

249
MCQhard

A fraud detection model is served on a Databricks Model Serving endpoint. You notice that predictions for the same input vector differ between two consecutive requests within seconds, and there is no feature store or external cache involved. The model was logged with a fixed random seed and deterministic inference code. Which action should you take first to diagnose the inconsistency?

A.Convert the model to a different flavor such as ONNX and re-register it, because the original flavor cannot guarantee deterministic inference.
B.Increase the endpoint's concurrency and enable autoscaling, since differing predictions indicate the replicas are running different code.
C.Inspect the endpoint's request logs to confirm the payloads are identical and check whether the model version serving traffic changed between requests.
D.Retrain the model with a different random seed and redeploy, because seed instability is the likely cause of varying predictions.
AnswerC

Non-determinism in a supposedly deterministic model most often comes from the serving layer, not the model math. Verifying that the request payloads are byte-identical rules out client-side variation, and checking whether the endpoint switched model versions or routed to a different version reveals the most common cause of divergent outputs. This is the cheapest, highest-signal first step before deeper debugging of the model artifact or environment.

Why this answer

When a deterministic model returns different outputs for the same input, the fastest path to the cause is to verify the two requests were truly identical and to confirm which model version handled each. A version transition during a rollout, or a payload difference such as field ordering or missing values, explains the symptom without any model retraining. Retraining, scaling, or changing flavors are premature and do not produce diagnostic evidence.

Exam trap

The trap here is jumping to model-level explanations like seeds or flavors, when serving-layer causes such as version routing or payload differences are far more likely.

250
MCQmedium

A machine learning team is transitioning from local notebooks to Databricks. They want to ensure their code is modular and reusable. Which THREE practices should they implement?

A.Package common utility code into Python wheels.
B.Keep all training, preprocessing, and evaluation code in a single notebook.
C.Integrate with Databricks Repos for version control using Git.
D.Use Delta Lake tables to ensure data versioning and consistency.
E.Hardcode all file paths and credentials in every notebook.
AnswerA, C, D

Creating Python wheels allows teams to share versioned, modular code across different projects and notebooks. This eliminates code duplication, simplifies dependency management, and enables unit testing of utility functions, which significantly improves the reliability and maintainability of machine learning pipelines in a collaborative multi-user development environment.

Why this answer

Transitioning to Databricks requires moving away from monolithic notebooks toward modular code structures. Using Delta Lake for data consistency, moving logic into Python wheels, and utilizing Repos for version control are the standard industry practices. These steps ensure that code is maintainable, testable, and capable of being integrated into automated CI/CD pipelines, which is the hallmark of mature MLOps practices within a Databricks workspace.

Exam trap

Candidates often include 'copy-pasting code across notebooks' or 'manual file versioning' as valid practices, failing to recognize that modularity requires wheels and formal version control via Repos.

251
MCQhard

When auditing an ML pipeline in Databricks for compliance and governance, which THREE of the following should be verified?

A.The Git commit hash associated with the specific training run.
B.The personal credentials of the data scientist who manually ran the training.
C.The lineage of the training data, including source table versions.
D.The list of authorized users who can transition models to 'Production'.
E.The raw text of every email sent between the data scientists.
AnswerA, C, D

Traceability starts with code. By linking every model to a specific Git commit, auditors can verify the exact code state that produced the model. This is fundamental for reproducibility and ensuring that no unauthorized or unreviewed changes were introduced into the model generation process.

Why this answer

Auditing requires proof of data lineage, code versioning, and access control. Verifying these components ensures that the organization can prove how a model was built, what data was used, who approved it, and that the code was properly reviewed. This level of traceability is non-negotiable for regulated industries and essential for enterprise-grade MLOps maturity.

Exam trap

Candidates often include irrelevant metrics like model accuracy or latency as audit requirements, failing to realize that compliance audits focus on traceability, lineage, and access control.

252
MCQmedium

A data scientist is training a scikit-learn model on Databricks and wants to capture the best hyperparameters found during a hyperparameter sweep. They are using MLflow Tracking with nested runs. Which approach correctly records the best parameters and metrics in the parent run?

A.Set the parent run's status to 'FINISHED' and then use `mlflow.log_artifact()` to attach a JSON file containing the best parameters.
B.Use `mlflow.log_params()` inside each child run and rely on MLflow to automatically propagate the best parameters to the parent run.
C.Log the best parameters and metrics directly in the parent run using `mlflow.log_params()` and `mlflow.log_metrics()` after the sweep completes.
D.Use `mlflow.start_run(nested=True)` for the parent run and `mlflow.start_run()` for child runs, then log the best parameters in the parent run.
AnswerC

This approach works because nested runs are children of the parent run; after all child runs finish, the parent run can log the aggregated best parameters and metrics using standard MLflow logging functions. It ensures the parent run contains a summary of the best configuration, which is useful for comparison and model selection.

Why this answer

In MLflow, nested runs allow you to organize hyperparameter sweeps. The parent run acts as a container, and child runs log individual trials. To capture the best parameters and metrics in the parent, you must explicitly log them after the sweep.

MLflow does not auto-propagate, so manual logging in the parent run is required.

Exam trap

The trap here is assuming that MLflow automatically aggregates or propagates the best results from nested runs to the parent run.

253
MCQhard

Your organization is implementing an MLOps strategy that requires strict model governance. You need to ensure that no model is deployed to production unless it has been tagged with 'validated=true' in the MLflow Model Registry. How can you enforce this policy within your CI/CD workflow?

A.Use Databricks workspace permissions to prevent unauthorized users from deploying models.
B.Add a validation script in the CI/CD pipeline that queries the model version tags.
C.Rely on the MLflow UI to visually verify the tag before clicking 'Deploy'.
D.Set the default stage of all new models to 'Production' via a global config.
AnswerB

A custom validation script acts as a gatekeeper. By using 'mlflow.tracking.MlflowClient()' to retrieve the model version and inspect its dictionary of tags, the pipeline can verify the 'validated=true' condition. If the condition is not met, the pipeline fails, preventing the deployment from occurring.

Why this answer

Implementing a policy-based deployment gate in your CI/CD pipeline ensures that governance requirements are met before code or models hit production. By querying the model's tags using the MLflow client before attempting a deployment, you create a hard stop for non-compliant models. This is a critical security and compliance practice, preventing unauthorized or untested models from being served to end-users or critical business processes.

254
MCQmedium

Which of the following describes the purpose of a 'Gold' table in the Medallion architecture within an MLOps pipeline?

A.It stores raw data ingested directly from external sources.
B.It contains validated, business-level data ready for ML model training.
C.It serves as a transient staging area for schema evolution.
D.It stores model artifacts and hyperparameters for versioning.
AnswerB

The Gold layer serves as the final consumption layer, providing reliable, high-quality data. In an MLOps context, this is the optimal source for training data, ensuring that the model is built on clean and consistent information, which minimizes the risk of garbage-in, garbage-out performance issues.

Why this answer

The Gold layer contains highly processed, business-level data that is ready for consumption by downstream ML models or analytical applications. By ensuring data is clean, aggregated, and validated at this stage, MLOps teams can rely on high-quality features for training, which directly improves model performance and reduces the complexity of the feature engineering step in the training pipeline.

Exam trap

Candidates confuse Gold tables with raw ingestion tables (Bronze) or intermediate cleaned tables (Silver), missing that Gold stores analytics-ready business data.

255
MCQmedium

A machine learning engineer has a model registered in Unity Catalog as prod.ml.iris_model. They need to deploy it to a real-time serving endpoint that automatically scales based on traffic and provides a REST API for predictions. The model's signature is logged. Which deployment method should they use?

A.Use MLflow's built-in serving command to start a local REST server on a cluster and expose it via a public URL.
B.Use Databricks Model Serving to create a new endpoint and select the model from Unity Catalog, specifying the model version or alias.
C.Create a Databricks job that runs a Python script to load the model and listen for HTTP requests on a driver node.
D.Export the model as a Docker image using MLflow and deploy it to a Kubernetes cluster managed outside Databricks.
AnswerB

Databricks Model Serving is the managed service for real-time inference. It integrates with Unity Catalog, allowing you to select a registered model by name and version or alias. The endpoint provides a REST API, auto-scales based on load, and handles the serving infrastructure, which matches the requirement for a scalable real-time endpoint.

Why this answer

Databricks Model Serving provides a fully managed, scalable solution for deploying models registered in Unity Catalog. It automatically creates a REST endpoint, handles scaling, and integrates with governance features. The other options either rely on manual infrastructure or are intended for development, not production real-time serving.

Exam trap

The trap here is assuming that any method that exposes a REST API, such as MLflow serve or a custom Flask app, is equivalent to a managed serving endpoint.

256
MCQeasy

A data scientist wants to compare multiple runs within the same MLflow experiment to identify the best-performing model based on a custom metric. What is the most efficient way to do this in the MLflow UI?

A.Export all runs to a CSV file using the MLflow CLI and then analyze the CSV in a separate tool.
B.Use the MLflow UI's experiment page to sort runs by the custom metric and compare them side by side.
C.Query the MLflow tracking server's REST API and write a custom script to parse and rank the runs.
D.Open each run individually and manually record the custom metric values in a spreadsheet.
AnswerB

The MLflow UI experiment page lists all runs and allows sorting by any logged metric, including custom ones. You can select multiple runs and compare their parameters, metrics, and artifacts side by side. This is the intended and most efficient way to identify the best run.

Why this answer

The MLflow UI experiment page is designed for comparing runs: it displays all runs in a table, supports sorting by any metric, and allows selecting multiple runs for side-by-side comparison of parameters, metrics, and artifacts. Manual recording, CSV export, or custom API scripts are less efficient and unnecessary for this common task.

Exam trap

The trap here is overlooking the built-in comparison features of the MLflow UI and instead reaching for manual or programmatic methods that add unnecessary work.

257
Multi-Selecthard

A team is using Databricks Feature Store to manage features for a real-time model served via Databricks Model Serving. They need to ensure that the online feature values used at inference time are consistent with the training data. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Log the model with the feature store's training set specification to enable automatic feature lookup at serving time.
B.Store all features in a single Delta table and query it directly from the model serving endpoint.
C.Use point-in-time lookups when creating the training dataset to avoid leaking future data.
D.Publish feature tables to an online store that is updated with the same pipeline that writes to the offline store.
E.Manually copy feature values from the offline store to the online store on a daily basis using a notebook.
AnswersC, D

Point-in-time lookups ensure that training examples only use feature values available at the time of the label event, mimicking real-time inference conditions. This prevents data leakage and aligns training with serving. Databricks Feature Store provides time-series lookups for this purpose, which is essential for temporal consistency.

Why this answer

Consistency between online and offline features is achieved by using the same pipeline to publish to both stores and by using point-in-time lookups during training. These practices ensure that the transformations and temporal alignment match between training and inference. Manual copying or direct Delta queries do not provide the required consistency or performance.

Exam trap

The trap here is thinking that logging the model with the training set specification alone guarantees consistency, when the online store must also be updated with the same pipeline.

258
MCQhard

A machine learning engineer is developing a custom PyFunc model that combines a scikit-learn preprocessing step and a TensorFlow model. They log the model with MLflow and specify a signature. When they attempt to serve the model using Databricks Model Serving, the endpoint returns errors about incompatible input types. The signature was inferred from a pandas DataFrame with integer columns, but the serving request sends JSON with floating-point numbers. Which modification to the model signature will resolve this issue?

A.Remove the signature entirely so the serving endpoint does not enforce input types.
B.Add a preprocessing step in the PyFunc model to cast all input columns to integers before passing to the TensorFlow model.
C.Change the signature's input data types from `long` to `double` for the affected columns.
D.Change the signature's input data types from `long` to `string` for the affected columns.
AnswerC

JSON numbers without a decimal point are inferred as integers, but if the signature expects `long` and the request sends floating-point numbers, the serving endpoint will reject them. Changing the signature to `double` allows both integer and floating-point numbers, as integers can be safely cast to double. This resolves the incompatibility and ensures the endpoint accepts the JSON input. This is the correct modification.

Why this answer

The signature inferred from integer columns specifies `long` data types, which do not accept floating-point numbers in JSON requests. Databricks Model Serving enforces the signature strictly. By changing the signature to `double`, the endpoint will accept both integer and floating-point numbers, as integers can be cast to double without loss.

This resolves the incompatibility without altering the model's logic. The other options either introduce incorrect types or remove validation, which is not advisable.

Exam trap

The trap here is assuming that removing the signature or casting inputs will fix the issue, when the correct solution is to update the signature to a compatible numeric type.

259
MCQmedium

A data scientist at your company has trained a scikit-learn model and logged it with MLflow using the default 'sklearn' flavor. The model must now be deployed to a Databricks Model Serving endpoint for real-time inference, and the endpoint must automatically reload the new version whenever the model's stage changes to 'Production' in the MLflow Model Registry. Which deployment approach should the team use?

A.Log the model to a new MLflow experiment and configure the endpoint to poll the experiment's artifact location for changes.
B.Package the model into a Docker image with a Flask app and deploy it to a Databricks cluster running as a job.
C.Export the model artifact to DBFS as a pickle file and create a Python script that loads the pickle at endpoint startup.
D.Register the model version in the MLflow Model Registry, then create the serving endpoint by specifying the model name and the 'Production' stage in the endpoint configuration.
AnswerD

Databricks Model Serving natively integrates with the MLflow Model Registry. When you configure an endpoint with a registered model name and a stage such as 'Production', the endpoint automatically serves the version currently in that stage and reloads when the stage transitions, satisfying the automatic-reload requirement without custom code.

Why this answer

The requirement is automatic reload when a model version enters the 'Production' stage. Databricks Model Serving supports this directly by referencing the registered model name and stage in the endpoint configuration. Other approaches either bypass the registry, require manual redeployment, or rely on mechanisms the platform does not provide, making them unsuitable for a managed real-time endpoint with stage-driven updates.

Exam trap

The trap here is assuming the endpoint must be recreated or manually updated each time a model version changes stage, when stage-based endpoint configuration already reloads automatically.

260
MCQmedium

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires custom post-processing logic and loading auxiliary tokenizer files alongside the serialized weights. Which approach provides the correct mechanism to package and serve this custom artifact?

A.Log the model as a standard torch.jit artifact without metadata and point the serving endpoint directly to the raw .pt file in Unity Catalog.
B.Create an external FastAPI application outside Databricks, wrap the model, and route requests via a custom reverse proxy.
C.Subclass mlflow.pyfunc.PythonModel, implement the load_context and predict methods, bundle the tokenizer files in artifacts, and log via mlflow.pyfunc.log_model.
D.Store the tokenizer files in a Delta table and configure the serving endpoint to query the table on every incoming inference request.
AnswerC

Subclassing mlflow.pyfunc.PythonModel satisfies the custom post-processing and auxiliary file constraints: load_context loads the bundled tokenizer artifacts at initialisation, while predict applies the bespoke logic around the PyTorch weights. Logging via mlflow.pyfunc.log_model registers the whole bundle, which Model Serving then deploys as a single custom artifact.

Why this answer

Custom PyTorch models requiring custom code and auxiliary files must be logged using MLflow's pyfunc flavor with custom artifact dependencies. This allows packaging artifacts and arbitrary Python code safely so that the serving infrastructure can instantiate the pyfunc wrapper, execute custom tokenization, and perform required post-processing cleanly during real-time inference.

Exam trap

Candidates mistakenly try to log raw PyTorch weight files directly without a pyfunc wrapper, failing to provide the required custom tokenization and post-processing logic inside the serving container.

261
MCQeasy

A machine learning engineer is training a model using scikit-learn on Databricks and wants to track the model's hyperparameters, metrics, and artifacts automatically without adding explicit logging calls. Which MLflow feature should they use?

A.mlflow.tracking.MlflowClient()
B.mlflow.set_experiment()
C.mlflow.log_artifact()
D.mlflow.sklearn.autolog()
AnswerD

mlflow.sklearn.autolog() automatically logs parameters, metrics, and the model artifact for scikit-learn models. It captures hyperparameters from the estimator, metrics from scoring functions, and saves the trained model. This eliminates the need for manual logging calls and is the standard way to track scikit-learn experiments in Databricks.

Why this answer

mlflow.sklearn.autolog() is the correct choice because it automatically logs scikit-learn model parameters, metrics, and artifacts. The other options are either low-level APIs that require manual logging or functions that only set experiment context or log individual artifacts, none of which provide automatic tracking.

Exam trap

The trap here is confusing MLflow tracking APIs like MlflowClient or log_artifact with autologging, which is the only feature that automatically captures scikit-learn training details.

262
MCQmedium

An ML engineer has registered a scikit-learn model in the Unity Catalog Model Registry and wants to serve it as a real-time endpoint using Databricks Model Serving. The model's MLflow signature expects a JSON payload with an array of records. The engineer creates a serving endpoint with a workload size of Medium and configures the served entity to use the latest model version. Which additional configuration is required to enable automatic payload logging to a Delta table for monitoring?

A.Enable inference tables on the endpoint by specifying a Unity Catalog table location.
B.Set the environment variable MLFLOW_ENABLE_SYSTEM_METRICS_LOGGING to true in the model's conda environment.
C.Attach a delivery log to the endpoint and specify a Delta table path.
D.Configure the model signature to include a 'log_payloads' parameter set to true.
AnswerA

Inference tables capture request and response payloads for served models. To enable them, you must configure the endpoint with a Unity Catalog table location where logs are written. This allows monitoring and debugging without altering the model artifact. The other options do not provide payload logging.

Why this answer

Automatic payload logging in Databricks Model Serving is achieved through inference tables, which require specifying a Unity Catalog table location when creating or updating the endpoint. This captures request and response data for monitoring and debugging. Other logging mechanisms like delivery logs or MLflow environment variables do not record inference payloads.

Exam trap

The trap here is confusing delivery logs with inference tables; delivery logs track endpoint build events, not request/response payloads.

263
MCQmedium

You are building a Databricks job that trains a model and registers it to the MLflow Model Registry. After registration, you need to automatically transition the model version to 'Staging' only if its validation accuracy exceeds 0.95. The job should fail if the accuracy is below this threshold. Which approach should you implement?

A.Use the MLflow Client to register the model, then use a conditional step in the job to call the transition_model_version_stage method if the accuracy metric meets the threshold.
B.Configure the MLflow Model Registry webhook to automatically transition the model version to 'Staging' when a new version is registered.
C.Use the Databricks REST API to update the model version stage after the job completes, relying on the job's success or failure to indicate accuracy.
D.Set the model version's stage to 'Staging' during registration using the stage parameter in mlflow.register_model, and rely on the training script to skip registration if accuracy is low.
AnswerA

The MLflow Client provides programmatic control over model versions, including transition_model_version_stage. By retrieving the run's metrics and conditionally transitioning only when accuracy exceeds 0.95, you enforce the business rule. If the condition fails, you can raise an exception to fail the job. This approach is flexible and integrates well with Databricks jobs.

Why this answer

To conditionally transition a model version based on a metric, you need to programmatically retrieve the metric and then call the MLflow Client's transition method. This allows you to enforce a threshold and fail the job if the condition is not met. Webhooks and REST API calls lack the ability to evaluate metrics before transitioning, and setting the stage during registration does not incorporate a conditional check.

Exam trap

The trap here is assuming that webhooks or registration-time stage settings can enforce a metric threshold, when they actually operate after or independently of metric evaluation.

264
MCQmedium

A data scientist is training a machine learning model on Databricks using MLflow. They need to track hyperparameter tuning experiments while ensuring that each iteration is uniquely identifiable and reproducible. Which feature should they use to group related runs within a single experiment?

A.MLflow Tags to append metadata for each run iteration.
B.MLflow Experiment IDs for every parameter variation.
C.MLflow Nested Runs via the 'parent_run_id' parameter.
D.MLflow Model Registry versions to track iterative progress.
AnswerC

Nested runs allow developers to logically group multiple experiment iterations under a single primary run. This structure is specifically designed for hyperparameter tuning, where a central parent run tracks the overall task while individual child runs capture specific configuration results, ensuring cleanliness and logical grouping for reporting.

Why this answer

MLflow nested runs are the industry-standard approach for hierarchical organization of hyperparameter tuning tasks. By creating a parent run for the overall optimization process and child runs for individual parameter sets, data scientists gain clear visibility into model performance metrics across the entire search space, simplifying comparative analysis and facilitating efficient model selection during the development lifecycle.

Exam trap

Candidates often try to use tags or manual naming conventions to group runs, missing that MLflow specifically provides the 'parent_run_id' feature to enable programmatic hierarchical organization of experiments.

265
Multi-Selecthard

An ML engineer is deploying a model to Databricks Model Serving and needs to enable automatic scaling based on traffic. The model has variable inference latency and the team wants to optimize cost while maintaining performance. Which TWO configurations are required to achieve this? (Choose two.)

Select 2 answers
A.Deploy the model using a GPU-enabled workload type to reduce latency.
B.Enable auto-scaling by specifying a target concurrency per replica.
C.Configure the minimum and maximum number of replicas for the endpoint.
D.Set the endpoint to use a dedicated cluster with autoscaling enabled.
E.Set the scale-to-zero option to true so the endpoint can shut down when idle.
AnswersB, C

Auto-scaling in Databricks Model Serving is driven by a target concurrency metric. You specify the desired number of concurrent requests per replica, and the system adjusts the replica count to maintain that target. This is essential for scaling based on traffic, as it directly ties replica count to request load.

Why this answer

To enable automatic scaling for a Databricks Model Serving endpoint, you must define the minimum and maximum replica counts and set a target concurrency per replica. These settings allow the serving infrastructure to adjust the number of replicas in response to traffic, ensuring performance while controlling cost. Other options either do not directly enable scaling or are not applicable to Model Serving.

Exam trap

The trap here is confusing cost-saving features like scale-to-zero with autoscaling, which requires replica bounds and a concurrency target.

266
MCQhard

A team has deployed a model to Databricks Model Serving and enabled inference tables. They notice that the inference table contains request and response payloads but no ground truth labels. They want to automatically join ground truth labels for monitoring. What should they do?

A.Use the MLflow Model Registry webhook to trigger a job that updates the inference table with ground truth labels.
B.Set the 'log_inputs' and 'log_outputs' flags to true in the endpoint configuration, which will automatically include ground truth labels.
C.Enable 'auto_capture_config' with 'ground_truth_column' set to the label column name in the inference table.
D.Configure a Databricks SQL query that joins the inference table with a ground truth Delta table and schedule it to refresh a monitoring metric.
AnswerD

Inference tables store request and response data, but ground truth labels must be provided separately. The standard approach is to create a Delta table containing ground truth labels and join it with the inference table using a unique request ID. A scheduled Databricks SQL query or a Lakeflow pipeline can perform this join and compute monitoring metrics.

Why this answer

Databricks Model Serving inference tables capture request and response payloads but do not include ground truth labels. To monitor model quality, you must join the inference table with a separate ground truth table. This is typically done by scheduling a query or pipeline that performs the join and computes metrics, which can then be used for model monitoring.

Exam trap

The trap here is believing that inference tables automatically capture ground truth labels or that a simple configuration flag can enable it, when in reality ground truth must be joined from an external source.

267
Multi-Selectmedium

A machine learning team is using Databricks Feature Store to manage features for their models. They want to ensure that the features used during training are consistent with those served in production. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Use `FeatureStoreClient.score_batch` to score data in batch mode, which automatically handles feature retrieval.
B.Manually copy feature values from the offline store to the online store before each scoring request.
C.Log the model with `FeatureStoreClient.log_model`, providing the feature lookups used during training.
D.Use `FeatureStoreClient.create_training_set` to build the training dataset, specifying the feature lookups.
E.Disable online store publishing to avoid data duplication and reduce costs.
AnswersC, D

Logging the model with `log_model` and including the feature lookups packages the model with metadata about the features it requires. When the model is served, Databricks can automatically retrieve the necessary features from the online store, ensuring consistency. This is a key practice to maintain parity between training and serving, as it links the model to the exact feature definitions.

Why this answer

To ensure consistency between training and serving with Databricks Feature Store, teams should use `create_training_set` to build training data from feature lookups and log models with `log_model` including those lookups. These practices embed the feature retrieval logic into the model, enabling automatic and consistent feature serving. Manual copying or disabling online publishing would introduce inconsistencies and are not recommended.

Exam trap

The trap here is thinking that any method that retrieves features, such as batch scoring, is sufficient for ensuring training-serving consistency, when the key is to use the training set creation and model logging with feature lookups.

268
MCQeasy

When logging a model to the MLflow Model Registry, what is the primary benefit of using a registered model name rather than just the model URI?

A.It automatically increases the model training speed during the next iteration.
B.It enables seamless staging and production management without code changes.
C.It forces the model to be saved in a specific proprietary cloud format.
D.It ensures that the model is automatically encrypted using hardware security modules.
AnswerB

The Model Registry allows you to transition versions between stages like 'Staging' and 'Production'. By pointing your application to a registered name, you can swap the underlying model version simply by updating its stage. This eliminates the need to modify application code or configuration files during model deployment cycles.

Why this answer

Registered model names provide a level of indirection that allows you to manage model lifecycle stages such as 'Staging', 'Production', and 'Archived'. By referencing a model name, you can update the underlying model version without requiring code changes in your inference applications. This promotes operational stability, as teams can transition models through an automated CI/CD pipeline while maintaining consistent endpoints for downstream consumers of the model.

Exam trap

Candidates think registered model names are purely organizational labels, missing their core role in enabling alias-based promotion without code changes.

269
MCQhard

Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?

A.The 'entity_version' is set to a legacy version that uses more expensive hardware.
B.The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.
C.The endpoint is using a GPU-accelerated instance by default for all Unity Catalog models.
D.Multiple versions of the model are being served simultaneously, doubling the cost.
AnswerB

When scale-to-zero is disabled, the system maintains the minimum number of provisioned instances (defaulting to 1). This provides the benefit of zero latency for the first request after an idle period but results in constant resource consumption and associated cloud costs.

Why this answer

The configuration shows that 'scale_to_zero_enabled' is set to false. This means that at least one instance of the model is running 24/7, regardless of whether any requests are being made. While this ensures there is never a cold start, it leads to continuous billing for the compute resources even during nights and weekends.

Exam trap

Candidates often assume endpoints automatically scale down to zero when idle, forgetting that scale-to-zero must be explicitly enabled to avoid continuous compute charges.

270
MCQhard

Which THREE of the following are primary responsibilities of an MLOps engineer when maintaining production ML models in Databricks?

A.Monitoring model performance and detecting data drift.
B.Writing the core machine learning research papers for the organization.
C.Establishing CI/CD pipelines to automate testing and deployment.
D.Configuring workspace security and access control for model artifacts.
E.Manually retraining every model daily regardless of performance metrics.
AnswerA, C, D

Drift detection is critical to maintaining model accuracy. As data distributions change, models can lose predictive power. MLOps engineers must implement monitoring solutions to identify these shifts early, allowing for timely retraining and redeployment, which protects the organization from deploying degraded models to production users.

Why this answer

An MLOps engineer's role is to ensure stability, performance, and compliance. This includes monitoring model performance to detect drift, managing the CI/CD pipeline for automated deployments, and ensuring security via robust access controls. By balancing these tasks, they ensure that the machine learning system remains reliable and valuable to the business over time, effectively bridging the gap between development and operations.

Exam trap

Candidates focus exclusively on model training metrics, neglecting essential MLOps operational tasks like drift monitoring, CI/CD, and security controls.

271
MCQmedium

Your team is using Databricks Feature Store to serve features to a real-time model deployed on Databricks Model Serving. A data scientist updates the feature computation logic for one of the features and publishes a new version to the online store. However, the model endpoint continues to return stale feature values for that feature. What is the most likely cause?

A.The model endpoint has not been restarted after the feature table update, so it is still using cached feature values.
B.The online store is not configured to automatically sync with the offline store, so the new feature values were never published.
C.The model was logged with a feature spec that references an older version of the feature table, so it continues to use the old feature values.
D.The model was logged without the feature lookup package, so it cannot query the online store.
AnswerC

When you log a model with Feature Store, the feature spec captures the exact feature table versions at that time. If you later publish a new version of a feature table, the existing model still references the version recorded in its feature spec. To use the updated feature, you must re-log the model with a new feature spec that points to the new version, then redeploy the endpoint.

Why this answer

The model's feature spec records the versions of the feature tables used during training and logging. When a new feature table version is published, existing models still reference the old version, leading to stale values at inference. To resolve this, the model must be re-logged with an updated feature spec that includes the new version, and the endpoint must be updated with the new model version.

Exam trap

The trap here is assuming that updating the feature table automatically updates all models that use it, when in fact each model's feature spec pins specific table versions.

272
Multi-Selecthard

A machine learning engineer is preparing a model for deployment using Databricks Model Serving. They need to ensure that the model's input schema is enforced and that the model can be served with a specific version. Which TWO actions should they perform? (Choose two.)

Select 2 answers
A.Enable autoscaling on the serving endpoint to handle varying load.
B.Log the model with an input example using mlflow.sklearn.log_model(..., input_example=...)
C.Specify the model version when creating or updating a serving endpoint.
D.Register the model in the MLflow Model Registry and transition it to Production.
E.Use mlflow.log_artifact() to save a JSON schema file alongside the model.
AnswersB, C

Providing an input example during model logging allows MLflow to infer and store the input schema. This schema is then used by Model Serving to validate incoming requests, ensuring that the data types and structure match what the model expects. This action directly helps enforce the input schema and prevents errors during serving. It is a recommended practice for production deployments.

Why this answer

Logging the model with an input example captures the input schema, which Model Serving uses to validate requests. Specifying the model version when creating the serving endpoint ensures the correct version is deployed. Together, these actions enforce schema and control versioning.

The other options manage lifecycle, log artifacts, or scale resources, but do not directly achieve the stated requirements.

Exam trap

The trap here is assuming that registering a model or logging a separate schema artifact automatically enforces input schema in Model Serving, when only the model signature does.

273
MCQmedium

A machine learning engineer is training a PyTorch model on a Databricks cluster and needs to distribute the training across multiple worker nodes. Which framework should be integrated natively within Databricks to handle this distributed deep learning workflow efficiently?

A.Databricks Feature Store Client
B.TorchDistributor
C.MLflow Model Registry
D.Spark MLlib Pipeline
AnswerB

TorchDistributor is Databricks' native utility for launching PyTorch distributed training, wrapping torch.distributed and handling node discovery, environment setup and inter-worker communication. It satisfies the stem's requirement to distribute training across worker nodes efficiently without manual cluster configuration.

Why this answer

TorchDistributor is the native Databricks library designed to launch distributed PyTorch training jobs seamlessly across cluster nodes using standard PyTorch native CLI commands and environment configurations. Understanding distributed training orchestration on Databricks is crucial for scaling deep learning pipelines on large datasets without manually managing cluster communication sockets.

Exam trap

Candidates often select generic distributed frameworks like Horovod or Dask, ignoring that Databricks provides a specific, native integration called TorchDistributor for PyTorch workflows.

274
Multi-Selecthard

A machine learning engineer is responsible for monitoring a production model deployed to Databricks Model Serving. The model predicts customer churn and is served via a REST endpoint. The engineer needs to detect data drift and model performance degradation over time. Which TWO actions should the engineer take to enable effective monitoring? (Choose two.)

Select 2 answers
A.Enable automatic model retraining triggered by any change in the input data schema.
B.Enable inference logging on the model serving endpoint to capture request and response payloads.
C.Use the model's training accuracy as a proxy for production performance and set up alerts based on that metric.
D.Configure the endpoint to use a smaller instance type to reduce cost, as monitoring does not require additional resources.
E.Schedule a Databricks job to periodically compute drift metrics by comparing logged inference data with the training dataset.
AnswersB, E

Inference logging captures the input features and model predictions for each request, which is essential for computing drift metrics and evaluating performance over time. Databricks Model Serving allows you to enable inference logging to a Delta table, where the data can be analyzed using Databricks SQL or notebooks. Without this logging, there is no historical record of production data to compare against training data or to compute accuracy metrics, making drift detection impossible.

Why this answer

Effective monitoring of a production model requires capturing inference data and analyzing it for drift and performance. Enabling inference logging provides the raw data, while scheduling a job to compute drift metrics against the training set enables detection of data drift. Together, these actions allow the engineer to identify when the model's input distribution diverges from training, which is a key signal of degradation.

Exam trap

The trap here is thinking that training accuracy or automatic retraining can substitute for actual production monitoring, when in fact you must capture and analyze live inference data to detect drift and degradation.

275
MCQhard

An ML engineer has a Databricks Model Serving endpoint that is currently serving a registered model version. A new model version is registered in Unity Catalog and must be rolled out to the endpoint without any downtime. Which approach should the engineer use to safely transition traffic to the new model version?

A.Modify the model version in Unity Catalog by overwriting the existing version with the new model artifacts.
B.Create a new endpoint with the new model version and manually update the client applications to point to the new endpoint URL.
C.Update the served entities of the endpoint to point to the new model version and rely on Databricks to perform a rolling update.
D.Delete the existing endpoint and recreate it with the new model version.
AnswerC

Databricks Model Serving supports updating the served entities of an existing endpoint to reference a new model version. The platform performs a zero-downtime rolling update, provisioning new capacity with the updated model before shifting traffic, so in-flight requests continue to be served by the old version until the new one is ready.

Why this answer

Updating the served entities of an existing endpoint is the supported method for zero-downtime model updates in Databricks Model Serving. The platform handles the rolling update by provisioning new resources with the updated model version and shifting traffic only when ready, so the endpoint remains available throughout the process.

Exam trap

The trap here is assuming that any change to the served model requires endpoint recreation or a new endpoint, when in fact updating the served entities of the existing endpoint triggers a zero-downtime rolling update.

276
MCQhard

Refer to the exhibit. You are loading a model from the registry. What does the 'models:/MyModel/1' URI specifically represent?

A.A direct pointer to the local file system path of the model artifact.
B.A unique identifier for the registered model name and its specific version.
C.An alias for the most recently trained model in the current experiment.
D.A reference to the raw source code used to train the model.
AnswerB

This URI is a versioned reference that maps to a specific, immutable artifact in the Model Registry. Using this format ensures that the loading process retrieves the correct, validated model version, which is critical for maintaining consistency and reliability in downstream inference applications and automated model serving pipelines.

Why this answer

The URI format 'models:/ModelName/Version' is the standard way to reference registered models within MLflow. It points to a specific, versioned artifact stored in the registry, ensuring that the inference code consistently uses the exact model state that was approved. This abstraction allows developers to change model versions without modifying the underlying inference application, promoting a stable production environment.

Exam trap

Candidates confuse run IDs with model URIs, often assuming the URI points to a temporary tracking run artifact rather than a versioned registry model.

277
MCQeasy

A data scientist wants to record the exact library dependencies and a code snapshot alongside a model so that a reviewer can later restore the same environment and reproduce the training run. They are logging with MLflow on Databricks. Which practice best satisfies this requirement?

A.Call `mlflow.log_artifact` on a generated `requirements.txt` and let autolog capture parameters only.
B.Log the model with `mlflow.sklearn.log_model` and pass pinned `pip_requirements` plus attach the source notebook as an artifact.
C.Log the model with `mlflow.sklearn.log_model` and rely on the cluster's installed libraries at load time.
D.Use `mlflow.autolog()` and set the experiment's artifact location to a Unity Catalog volume.
AnswerB

Pinning `pip_requirements` writes an exact dependency manifest into the model's MLmodel metadata, so restoration tooling can rebuild the environment. Attaching the source notebook preserves the code snapshot. Together these give the reviewer both the environment and the code needed to reproduce the training run faithfully, which is precisely what was requested.

Why this answer

Reproducibility requires capturing two things: the environment and the code. Pinning `pip_requirements` when logging the model embeds a version-exact dependency manifest that MLflow restoration consumes. Attaching the source notebook records the training logic as it existed at that moment.

Storing artifacts elsewhere or relying on autolog alone leaves one of these gaps, so the run cannot be faithfully recreated later.

Exam trap

The trap here is believing that autolog or artifact storage location alone captures the dependency environment, when the environment must be explicitly pinned.

278
MCQhard

A machine learning engineer is deploying a model to Databricks Model Serving and wants to implement a blue-green deployment strategy. They have registered two model versions in Unity Catalog: version 1 (current production) and version 2 (new candidate). They want to route 10% of traffic to version 2 for testing while keeping 90% on version 1. Which feature should they use to achieve this?

A.Create two separate endpoints and use a load balancer to distribute traffic based on weights.
B.Use the model registry's stage transitions to mark version 2 as 'Staging' and version 1 as 'Production', then enable automatic traffic mirroring.
C.Configure the endpoint with two served entities and use traffic splitting percentages.
D.Deploy version 2 as a separate endpoint and use Unity Catalog aliases to switch traffic instantly.
AnswerC

Databricks Model Serving supports serving multiple model versions within a single endpoint by defining multiple served entities. Each served entity can be assigned a percentage of traffic. By setting version 1 to 90% and version 2 to 10%, the engineer achieves a blue-green or canary deployment. This allows safe testing of the new version.

Why this answer

Databricks Model Serving allows multiple served entities per endpoint, each with a traffic percentage. This enables canary or blue-green deployments by routing a portion of traffic to a new model version while the rest goes to the stable version. Other methods like separate endpoints or aliases do not provide built-in percentage-based splitting.

Exam trap

The trap here is confusing model registry aliases or stages with traffic splitting; aliases switch all traffic, while traffic splitting within an endpoint allows gradual rollout.

279
MCQmedium

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model expects a single feature vector of 10 float values per request. The endpoint must return predictions in under 100 ms. Which approach should the engineer use to minimize per-request overhead?

A.Log the model with a signature that specifies a tensor input and enable the endpoint's request batching.
B.Enable autoscaling on the endpoint and send requests with a batch size of 1.
C.Use MLflow's pyfunc flavor with a custom predict method that processes one row at a time.
D.Log the model using the native scikit-learn flavor and deploy it with a workload size that provides sufficient CPU.
AnswerD

The native scikit-learn flavor allows the model server to load and invoke the model directly without an extra Python wrapper, reducing per-request overhead. Choosing an appropriate workload size ensures enough CPU resources to handle the model's computation quickly. This combination minimizes framework overhead and provides the necessary compute, making it the best choice to meet the sub-100 ms latency requirement for single-vector requests.

Why this answer

For low-latency single-vector inference, minimizing framework overhead is critical. The native scikit-learn flavor avoids the extra Python layer that pyfunc introduces, and a suitable workload size provides the CPU needed to execute the model quickly. Autoscaling and request batching address throughput rather than per-request latency, so they do not help meet the strict latency target.

Exam trap

The trap here is assuming that autoscaling or request batching reduces the latency of a single small request, when they primarily improve throughput under load.

280
MCQmedium

A data scientist is developing a model on Databricks and wants to use MLflow to compare multiple runs. They need to quickly identify the run with the lowest validation loss. Which MLflow UI feature allows them to sort and filter runs based on metrics?

A.The runs table in the MLflow experiment UI, where you can sort by metrics and apply filters.
B.The artifacts tab of a run, where you can view logged metric files and manually sort them.
C.The notebook's output cell, where you can print a DataFrame of runs and sort it.
D.The model registry page, where you can compare model versions by their metrics.
AnswerA

The MLflow experiment UI provides a runs table that displays all runs with their parameters, metrics, and tags. You can sort the table by any metric column, such as validation loss, and apply filters to narrow down runs. This makes it easy to identify the best run based on specific criteria.

Why this answer

The MLflow experiment UI includes a runs table that lists all runs with their metrics. You can sort by any metric, such as validation loss, to find the lowest value. Filters can also be applied to focus on specific runs.

This is the standard way to compare runs visually.

Exam trap

The trap here is thinking that the model registry or artifacts tab provides run comparison capabilities, when actually the experiment UI's runs table is designed for that purpose.

281
MCQmedium

Your ML pipeline requires a complex environment with specific C++ dependencies. What is the recommended way to manage this in Databricks?

A.Install dependencies in a setup script at runtime.
B.Use Databricks Container Services with a custom image.
C.Include dependencies in the MLflow model artifact.
D.Create a shared library JAR file.
AnswerB

Custom Docker images allow you to pre-install complex C++ dependencies, system libraries, and specific compilers that aren't available in the standard Databricks runtime. This ensures that every cluster node starts with an identical, fully configured environment, providing the stability and reproducibility required for complex MLOps pipelines.

Why this answer

Databricks Container Services allows users to define custom Docker images as the execution environment. This is the only way to handle complex system-level dependencies like C++ libraries that cannot be installed via standard pip or conda commands. By using custom containers, teams ensure that the training and inference environments are identical and include all necessary system-level components.

Exam trap

Candidates often choose standard cluster libraries or init scripts, forgetting that complex system-level C++ dependencies require custom Docker images via Databricks Container Services.

282
MCQeasy

Which component of Databricks ML is best suited for tracking model hyperparameters, metrics, and code versions during the experimentation phase of the ML lifecycle?

A.Databricks SQL
B.MLflow Tracking
C.Delta Lake
D.Unity Catalog
AnswerB

MLflow Tracking provides the API and UI to record experiments. It allows developers to log parameters, performance metrics, and artifacts like serialized models and plots, which are essential for comparing runs and selecting the optimal model for production deployment.

Why this answer

MLflow Tracking is the dedicated component designed to record every experiment's parameters, metrics, and code state. By capturing these elements, data scientists can compare different iterations of a model to identify the best-performing configuration. This visibility is vital in MLOps, as it creates an audit trail of how a final model was derived, supporting reproducibility and informed decision-making before promoting a model to the registry.

283
MCQmedium

An ML engineer is training an XGBoost model on Databricks and wants to leverage hyperparameter tuning using Hyperopt while automatically logging all trial parameters, metrics, and models to MLflow. Which built-in MLflow function should be used to achieve this automatic integration?

A.mlflow.xgboost.autolog()
B.mlflow.spark.autolog()
C.mlflow.sklearn.autolog()
D.mlflow.register_model()
AnswerA

`mlflow.xgboost.autolog()` hooks directly into XGBoost's training callbacks, capturing each Hyperopt trial's parameters, metrics and resulting model into MLflow runs without manual logging code. This satisfies the stem's requirement for automatic trial logging during hyperparameter tuning, whereas generic `mlflow.autolog()` covers fewer framework-specific details.

Why this answer

mlflow.xgboost.autolog() automatically logs parameters, metrics, and trained artifacts during XGBoost training routines without requiring manual logging statements inside the training loop. This function streamlines model development workflows by ensuring complete lineage tracking and experiment reproducibility across extensive hyperparameter tuning sweeps.

Exam trap

Candidates often select manual logging methods or generic MLflow functions instead of the specific library-native autologging function, failing to realize that autologging is the most efficient way to capture Hyperopt trials.

284
MCQhard

You are using Databricks Feature Store to manage features for a model. A feature table is updated daily with new data. Your model training job reads from the feature table and logs the model with MLflow. To ensure that the model in production uses the correct feature values at inference time, what must you do when logging the model?

A.Log the model with a static copy of the feature values used during training.
B.Log the model with the feature table's primary key and timestamp columns only.
C.Log the model with the feature store lookup tables and specify the feature table names.
D.Log the model without any feature store metadata and manually join the feature table in the scoring job.
AnswerC

When logging a model that uses Databricks Feature Store, you must include the feature store metadata by specifying the feature table names. This allows the model to automatically look up the latest feature values at inference time. MLflow integrates with Feature Store to package this information, ensuring consistency between training and serving.

Why this answer

To ensure the model uses the correct feature values at inference time, you must log the model with the feature store lookup tables and specify the feature table names. This enables automatic feature retrieval, maintaining consistency between training and serving and ensuring that daily updates are reflected in predictions.

Exam trap

The trap here is assuming that logging the model alone is sufficient, or that static features are acceptable, when in fact Feature Store integration requires explicit metadata to enable automatic lookups.

285
MCQhard

A financial institution deploys a credit scoring model using Databricks Model Serving. The model must log all incoming requests and outgoing responses to a Delta table for auditing. The ML engineer needs to enable this logging with minimal performance impact. Which solution should they implement?

A.Enable inference logging on the serving endpoint, specifying a Delta table location for logs.
B.Implement a custom logging wrapper in the model's predict method that writes each request to a Delta table.
C.Use a Databricks job to periodically query the endpoint's access logs and write them to a Delta table.
D.Configure the model to send logs to an external Kafka topic, then use a Databricks job to ingest from Kafka into Delta Lake.
AnswerA

Databricks Model Serving provides built-in inference logging that captures request and response payloads and writes them to a Delta table. This feature is designed for auditing and monitoring, and it operates asynchronously to minimize impact on inference latency. Configuring the endpoint with a Delta table path enables this logging.

Why this answer

Databricks Model Serving includes a built-in inference logging feature that asynchronously logs request and response data to a Delta table. This is the most efficient and least intrusive method. Custom wrappers or external systems add latency and complexity.

Access logs do not contain payloads. Therefore, enabling inference logging directly is the correct solution.

Exam trap

The trap here is assuming that access logs contain full payloads or that custom logging is necessary, when Databricks provides a native asynchronous logging feature.

286
MCQmedium

Your team uses MLflow Projects to package training code. A colleague runs the project on a Databricks cluster and it fails with a dependency conflict because the cluster has an older version of a library than the project's conda environment specifies. What is the most reliable way to ensure the project uses its declared dependencies without modifying the shared cluster?

A.Run the MLflow project with the --experiment-name flag to isolate dependencies per experiment.
B.Use the MLflow Projects CLI with the --conda flag or run via mlflow run, which creates an isolated conda environment from the project's conda.yaml.
C.Install the required library version globally on the cluster and restart it before running the project.
D.Convert the project to a Databricks notebook and use %pip install at the top of the notebook.
AnswerB

MLflow Projects supports conda environments defined in conda.yaml. Running with mlflow run and conda enabled creates an isolated environment that installs the declared dependencies, avoiding conflicts with the cluster's preinstalled libraries. This is the intended mechanism for reproducible dependency management.

Why this answer

MLflow Projects are designed to encapsulate dependencies in a conda.yaml, and running them with conda enabled creates an isolated environment that installs the specified versions. This avoids conflicts with shared cluster libraries and preserves reproducibility. Flags controlling experiment names, global installs, or notebook conversions do not provide the same isolation or reliability.

Exam trap

The trap here is assuming that any MLflow CLI flag or notebook-level pip install provides the same dependency isolation as the project's conda environment.

287
Multi-Selectmedium

You are designing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically retrain the model when new data arrives, validate it, and deploy it to a serving endpoint if it passes. Which two components are essential to achieve this? (Choose two.)

Select 2 answers
A.MLflow Tracking to log parameters, metrics, and artifacts for each run.
B.A Databricks Repos integration to version control notebooks and manage code changes.
C.A Databricks Model Serving endpoint that can be updated with the new model version via the MLflow API.
D.A feature store that automatically ingests new data and updates features in real time.
E.A Databricks Job that runs the training and validation notebooks on a schedule or trigger.
AnswersC, E

A Model Serving endpoint is essential for deploying the model and serving predictions. The pipeline must update the endpoint with the new model version after validation passes. This is typically done via the MLflow API or Databricks CLI. Without a serving endpoint, the model cannot be consumed in real time, so this is a critical component of the deployment step.

Why this answer

A Databricks Job provides the orchestration to run training and validation when new data arrives, and a Model Serving endpoint is required to deploy the validated model for real-time predictions. Together, they form the backbone of an automated retraining and deployment pipeline. MLflow Tracking, feature stores, and Repos are valuable but not strictly essential for the pipeline to function.

Exam trap

The trap here is overemphasizing supporting tools like MLflow Tracking or feature stores, which are helpful but not essential for the core automation and deployment loop.

288
Multi-Selecthard

You are auditing a Databricks environment to ensure compliance. Which TWO actions ensure the highest level of model lineage and reproducibility for models registered in MLflow?

Select 2 answers
A.Include the Git hash in the log_model metadata.
B.Use random seeds for all ML libraries.
C.Log the environment configuration (pip/conda) with the model.
D.Store all training data as CSV files in the workspace.
E.Manually copy artifacts to a shared folder.
AnswersA, C

Recording the Git hash during logging directly links the model artifact to the exact codebase used for training. This enables developers to checkout the precise state of the repository at the time of training, ensuring that lineage is preserved and debugging becomes straightforward.

Why this answer

Linking models to specific Git commits and tracking the environment dependencies via Conda or pip requirements files are essential for reproducibility. These steps ensure that when a model is redeployed, the exact code version and environment can be recreated, fulfilling the core MLOps requirement of auditability and consistent model behavior across different computing environments.

Exam trap

Candidates often focus on model accuracy metrics, ignoring that compliance and reproducibility require linking the model to specific code versions (Git) and exact dependency environments (Conda/pip).

289
MCQmedium

A team is developing a model on Databricks and wants to run an automated hyperparameter search over a scikit-learn pipeline. They need to try many parameter combinations in parallel across cluster workers while keeping every trial's parameters and metrics in MLflow. Which Databricks capability should they use to orchestrate the search?

A.Hyperopt with the `SparkTrials` backend passed to `fmin`.
B.A single-node grid search loop using `sklearn.model_selection.GridSearchCV` with `n_jobs=-1`.
C.Delta Live Tables pipelines with a `foreach` flow for each parameter set.
D.MLflow Projects with a `Multirun` entry point targeting the driver.
AnswerA

`SparkTrials` distributes each hyperparameter trial as a Spark job task across the cluster's executors, so many configurations run concurrently instead of sequentially. Each trial still logs parameters and metrics to the active MLflow run through the MLflow integration, giving the team both parallel search and centralized experiment tracking in one workflow.

Why this answer

The distributed hyperparameter search capability on Databricks is delivered by the Hyperopt integration with the `SparkTrials` backend. Passing `SparkTrials` to `fmin` turns each evaluation into a Spark task that runs on executors, allowing many configurations to be explored at once. Because the integration reports each trial to MLflow, the team gets parallel throughput and a complete, queryable record of every parameter set and its resulting metric.

Exam trap

The trap here is choosing core-based parallelism such as n_jobs=-1, which scales only within one machine and leaves cluster executors unused.

290
MCQhard

Your organization requires a feature store strategy that supports low-latency point-in-time lookups for online inference. Which implementation approach best addresses this?

A.Using a temporary Spark SQL view for features.
B.Utilizing the Databricks Feature Store with online store integration.
C.Caching all feature data in local cluster memory.
D.Performing manual joins in inference notebooks.
AnswerB

The Databricks Feature Store is specifically designed to manage feature pipelines, ensuring data consistency between offline training and online serving. Integrating with a high-performance online store allows for low-latency lookups, while the built-in point-in-time join logic prevents training-serving skew and data leakage.

Why this answer

Databricks Feature Store enables point-in-time joins to avoid data leakage during training and supports publishing to low-latency online stores (e.g., DynamoDB or Cosmos DB). By using a centralized feature store, teams ensure consistency between training and serving. Point-in-time lookups are crucial to ensure that features used for predictions reflect exactly what was known at that specific moment, preventing bias.

Exam trap

Candidates often confuse batch feature tables with online store deployments, failing to recognize that low-latency online inference requires dedicated online store integration.

291
MCQmedium

When deploying a model as a real-time REST endpoint on Databricks, how can you ensure the infrastructure scales automatically to handle increased request traffic?

A.Configure the cluster to use a fixed number of workers in the cluster settings.
B.Deploy the model using Databricks Model Serving endpoints.
C.Manually add instances to the inference cluster using the Databricks API before peak times.
D.Enable 'Auto-terminate' on the inference cluster configuration.
AnswerB

Databricks Model Serving endpoints are designed specifically for high-performance, real-time inference. They automatically handle infrastructure provisioning and scaling, adjusting the number of active model instances based on the current load. This abstracts away the complexity of cluster management, ensuring reliable and efficient model serving in production.

Why this answer

Databricks Model Serving provides serverless, auto-scaling inference endpoints. By default, it manages the underlying infrastructure, scaling the number of replicas based on real-time request volume. This is a key MLOps requirement to ensure high availability and performance without manual intervention, allowing teams to handle unpredictable traffic spikes without over-provisioning and incurring unnecessary costs during low-traffic periods.

Exam trap

Candidates often choose manual cluster configuration or custom scaling scripts, failing to realize that Model Serving endpoints provide native, serverless auto-scaling without needing manual infrastructure overhead.

292
MCQmedium

A data scientist is using MLflow to track experiments on Databricks. They notice that some runs are missing the model artifact even though they called mlflow.sklearn.log_model(). What is the most likely cause?

A.The model artifact was overwritten by a subsequent run with the same name.
B.The model artifact was logged outside of an active MLflow run.
C.The model artifact was too large and exceeded the maximum artifact size.
D.The model artifact was logged to a different experiment than the one being viewed.
AnswerB

MLflow requires an active run to log artifacts. If mlflow.sklearn.log_model() is called without an active run context, the artifact is not associated with any run and may be lost or logged to a default location. This is a common mistake when the logging call is placed outside a with mlflow.start_run() block. Ensuring the call is within an active run resolves the issue.

Why this answer

MLflow only logs artifacts when there is an active run. Calling mlflow.sklearn.log_model() outside a run context results in the artifact not being associated with any run, so it appears missing. Other options like size limits or overwrites are less likely.

Ensuring the logging call is inside a with mlflow.start_run() block or after mlflow.start_run() will correctly attach the model artifact to the run.

Exam trap

The trap here is overlooking the need for an active run context when logging artifacts, assuming the function works independently.

293
Multi-Selecthard

When designing a model training pipeline, which TWO features of Unity Catalog best support compliance and model governance?

Select 2 answers
A.Column-level access control to restrict sensitive data visibility during feature engineering.
B.Automated hyperparameter grid search for all registered datasets.
C.System-wide data lineage tracking that maps from raw data to the final registered model.
D.Built-in model serving endpoints for real-time inference.
E.Automatic translation of SQL queries into Python code.
AnswersA, C

Column-level access control allows organizations to define granular permissions, ensuring that data scientists only see the columns necessary for their specific tasks. This is vital for compliance with data privacy regulations like GDPR or HIPAA, as it prevents the accidental exposure of sensitive PII during the feature development phase.

Why this answer

Unity Catalog acts as a centralized governance layer for all data and AI assets. By providing fine-grained access control and end-to-end lineage, it ensures that only authorized users can access sensitive training data and that the origin of every model can be traced back to the original source data, which is essential for meeting regulatory requirements and maintaining corporate security standards.

Exam trap

Candidates often confuse workspace-level permissions or generic cloud storage policies with Unity Catalog's specific fine-grained governance capabilities like column-level access control and system-wide data lineage tracking.

294
MCQmedium

A data scientist is training a deep learning model on Databricks. They observe that the training process is significantly slower than expected. Upon inspection, they find that data loading from DBFS is the bottleneck. What is the most effective way to improve data loading speed for deep learning training on Databricks?

A.Move all data to local disk on the driver node.
B.Convert the dataset into a single large CSV file.
C.Use the Petastorm library to read data directly from Delta tables.
D.Increase the learning rate to reduce training time.
AnswerC

Petastorm enables efficient streaming of data from Apache Parquet files (such as those in Delta tables) directly into deep learning frameworks like TensorFlow and PyTorch. It is designed to maximize throughput by leveraging parallel reads across the Databricks cluster, which is critical for overcoming I/O bottlenecks during model training.

Why this answer

Using Petastorm or the standard 'tf.data' dataset API with Delta Lake significantly improves throughput for deep learning models. By converting data into a highly efficient, partitioned format that can be streamed directly to GPU memory, developers bypass the latency overhead associated with reading small files from DBFS, enabling the model to utilize the full processing power of the allocated cluster.

Exam trap

Test-takers frequently select generic cluster scaling options or standard Spark configurations, forgetting that deep learning specifically benefits from Petastorm or tf.data reading directly from Delta tables.

295
MCQeasy

An ML engineer needs to deploy a model to Databricks Model Serving. The model was logged with MLflow and registered in Unity Catalog. The engineer wants to ensure that only the latest version of the model is served and that the endpoint can be updated without downtime. Which approach should they use?

A.Delete the existing endpoint and create a new one with the same name and configuration, pointing to the new model version.
B.Use MLflow's transition request to move the model to Production stage, which automatically updates the serving endpoint.
C.Create a new endpoint for each model version and update the client application to point to the new endpoint URL.
D.Update the existing endpoint to serve the new model version using the Databricks UI or REST API, which performs a rolling update.
AnswerD

Databricks Model Serving allows you to update an endpoint to a new model version without downtime. The service performs a rolling update, gradually shifting traffic to the new version while maintaining availability. This ensures that the latest version is served and clients continue to use the same endpoint URL.

Why this answer

Updating an existing endpoint to a new model version via the UI or REST API triggers a rolling update, ensuring zero downtime and keeping the same endpoint URL. This is the standard method for deploying a new version without disrupting clients.

Exam trap

The trap here is assuming that MLflow stages or aliases automatically update serving endpoints, or that recreating an endpoint is necessary for updates.

296
MCQhard

An ML engineer is deploying a model to Databricks Model Serving and wants to implement A/B testing between two model versions. The engineer needs to route a percentage of traffic to each version and collect performance metrics. Which feature of Databricks Model Serving should the engineer use?

A.Deploy two separate endpoints and use an external load balancer to distribute traffic between them.
B.Enable the 'Canary' deployment option in the endpoint configuration to automatically split traffic.
C.Create a single endpoint with multiple model versions and configure traffic splitting between them.
D.Use MLflow's model registry to assign a stage to each model version and route traffic based on the stage.
AnswerC

Databricks Model Serving supports serving multiple model versions on a single endpoint with traffic splitting. You can specify the percentage of traffic routed to each version, enabling A/B testing. This is the built-in feature for such scenarios, allowing you to compare performance and metrics.

Why this answer

To perform A/B testing between model versions, the engineer should create a single Model Serving endpoint that serves multiple model versions with traffic splitting. This allows a specified percentage of requests to be routed to each version, and Databricks provides metrics for each version, facilitating performance comparison. This native feature simplifies A/B testing.

Exam trap

The trap here is thinking that MLflow stages or separate endpoints with external load balancers are needed for A/B testing, when Databricks Model Serving provides built-in traffic splitting on a single endpoint.

297
MCQmedium

A team runs a weekly retraining job that produces a new model version in Unity Catalog. Their production endpoint is currently serving version 4. They want to promote version 5 with zero downtime and the ability to roll back instantly if error rates rise. Which approach best meets these requirements?

A.Delete version 4 from the registry, deploy version 5, and recreate the endpoint so it picks up the new version automatically.
B.Configure the serving endpoint to serve version 5 and rely on the platform's rolling update so traffic shifts gradually while the previous configuration remains available for rollback.
C.Set the endpoint's served entities to version 5 with a traffic percentage of zero, then raise traffic to one hundred percent after validation.
D.Create a second endpoint for version 5, run both endpoints in parallel indefinitely, and split traffic manually at the load balancer.
AnswerB

Databricks Model Serving performs rolling updates that keep the endpoint available while new model versions warm up, and the prior configuration can be restored if problems appear. Pointing the endpoint at the new version and letting the rolling update proceed gives zero-downtime promotion with a fast rollback path, which matches both stated requirements without extra tooling.

Why this answer

Model Serving supports rolling updates that keep the endpoint live while a new model version loads, and the previous configuration remains restorable. Repointing the endpoint to the newer version and letting the rolling update run achieves zero-downtime promotion with a straightforward rollback. Deleting the old version or duplicating endpoints introduces risk or cost without improving the outcome.

Exam trap

The trap here is conflating a rolling update on a single endpoint with running duplicate endpoints or deleting the prior version.

298
MCQhard

A team has deployed a model to Databricks Model Serving and wants to enable autoscaling to handle variable traffic. They configure the endpoint with scale_to_zero_enabled set to true and a min_provisioned_concurrency of 0. After deployment, they notice that the endpoint takes several seconds to respond to the first request after a period of inactivity. What is the cause of this latency?

A.The model is being reloaded from Unity Catalog on each request, causing cold-start latency.
B.The endpoint is using a GPU workload, and GPU initialization takes several seconds after each idle period.
C.The model's Python environment is being reinstalled on each request due to missing dependencies in the container image.
D.The endpoint is scaling from zero, which requires provisioning resources and loading the model before serving the request.
AnswerD

When scale_to_zero_enabled is true and min_provisioned_concurrency is 0, the endpoint scales down to zero replicas during inactivity. The first request after idle time triggers a cold start: the system must provision compute resources and load the model into memory. This provisioning and loading time causes the several-second latency observed, which is inherent to scale-to-zero behavior.

Why this answer

Enabling scale_to_zero with zero minimum concurrency allows the endpoint to shut down completely when idle, reducing cost. However, the next request must wait for compute resources to be provisioned and the model to be loaded, resulting in cold-start latency. To avoid this, set a min_provisioned_concurrency greater than zero or disable scale_to_zero, keeping at least one replica warm.

Exam trap

The trap here is attributing cold-start latency to model artifact retrieval or environment setup instead of the scale-from-zero provisioning process.

299
MCQmedium

Refer to the exhibit. What happens to these logged metrics in MLflow when the training run completes?

A.MLflow overwrites the previous value with the latest value.
B.MLflow keeps all values, enabling the visualization of metrics over time.
C.The run will throw an error because the metric key is not unique.
D.MLflow averages the values and stores the mean.
AnswerB

Storing multiple values for a single metric key creates a sequence that MLflow can plot. This is vital for analyzing the model's convergence behavior, allowing data scientists to identify potential issues like overfitting or high variance early, and helping to fine-tune the training process effectively during the experiment cycle.

Why this answer

MLflow is designed to track metrics over time, which is crucial for monitoring model convergence during training. By logging the same metric key multiple times, MLflow stores each value as a point in a time series. This allows developers to visualize the training progress and identify when the model stops improving, which is critical for implementing early stopping and optimizing training efficiency in deep learning.

Exam trap

Candidates frequently assume MLflow overwrites previous metric values when updated, not realizing that MLflow natively supports time-series logging for every call of log_metric during a run.

300
MCQmedium

A machine learning engineer needs to track hyperparameter tuning experiments in Databricks using MLflow. Which approach best ensures that model training runs are associated with the correct code version and environment settings?

A.Manually log the git commit hash as a parameter in every mlflow.log_param call within the training script.
B.Utilize mlflow.set_tracking_uri with a local file system path for all distributed training nodes.
C.Run the training notebook from a Databricks Repo and use the mlflow.tracking.fluent API to track experiments.
D.Hardcode the environment configuration inside the model training loop using environment variables.
AnswerC

Databricks automatically captures the git context, including the branch and commit hash, when executing notebooks within a Repo. This integration ensures that experiment metadata is automatically enriched with source control information, providing a verifiable link between the model development process and the specific code repository state.

Why this answer

Integrating MLflow with Git projects via Databricks Repos allows for automatic logging of the git commit hash. This practice is crucial for reproducibility, as it enables data scientists to map specific model performance metrics back to the exact codebase state used during development, ensuring auditability and consistency across development, staging, and production environments.

Exam trap

Candidates often choose manual file uploads or standard local scripts, overlooking how Databricks Repos automatically integrates with Git to track exact code versions for experiment reproducibility.

Page 3

Page 4 of 4

All pages

Practice Databricks-ML-Pro by domain

Target a specific domain to shore up weak areas.

See all domains with question counts →