Courseiva

CCNA ML Ops Questions

75 of 130 questions · Page 1/2 · ML Ops · Answers revealed

1
MCQhard

Your organization wants to monitor production models for drift. Which Databricks service should be used to detect changes in the input data distribution compared to the training data?

A.MLflow Model Registry
B.Databricks Lakehouse Monitoring
C.Delta Live Tables (DLT)
D.Databricks SQL Alerting
AnswerB

Lakehouse Monitoring is specifically built to compute drift metrics by comparing production data with training baselines. It provides automated alerts and dashboards, enabling teams to proactively identify when input feature distributions shift, which is a primary cause of model performance degradation in production environments.

Why this answer

Databricks Lakehouse Monitoring is the native solution for tracking data quality and drift. By analyzing incoming data against a baseline established during training, it identifies statistical deviations. This is critical for MLOps because detecting drift early allows data scientists to trigger retraining or investigate data pipeline issues, ensuring model performance remains consistent over time despite changing real-world data patterns.

Exam trap

Test-takers frequently look for manual logging solutions, missing the native Databricks Lakehouse Monitoring service designed explicitly for automated drift detection.

2
MCQeasy

A data scientist is using MLflow Tracking to log experiments for a model that predicts customer lifetime value. They want to compare runs across multiple experiments and identify the best performing run based on a custom metric called 'rmse'. Which MLflow feature should they use?

A.MLflow Models
B.MLflow Model Registry
C.MLflow Projects
D.MLflow Tracking UI and API
AnswerD

The MLflow Tracking UI allows users to visualize and compare runs across experiments, including metrics like 'rmse'. The API (e.g., MlflowClient.search_runs) enables programmatic querying and filtering based on metrics. This is the correct tool for comparing runs and identifying the best one.

Why this answer

MLflow Tracking provides both a UI and an API to log, query, and compare runs. The UI allows visual comparison of metrics like 'rmse' across experiments, and the API supports programmatic search and filtering. This makes it the appropriate feature for identifying the best performing run based on a custom metric.

Exam trap

The trap here is confusing MLflow Tracking with MLflow Model Registry, which manages model lifecycle but not run comparison.

3
MCQmedium

Refer to the exhibit. An engineer is configuring a serving endpoint. Based on the configuration provided, what is the impact of the 'auto_scale' flag?

A.It scales the number of model versions deployed in the registry.
B.It allows the serving endpoint to adjust worker count based on request load.
C.It automatically re-trains the model when request volume drops.
D.It forces the endpoint to only use 10 workers permanently.
AnswerB

Auto-scaling is specifically designed to manage infrastructure resources dynamically. By adjusting the number of workers based on demand, the endpoint remains responsive while avoiding the unnecessary cost of provisioning for maximum load at all times, making it a best practice for production deployments.

Why this answer

The 'auto_scale' flag allows the Databricks Model Serving endpoint to dynamically adjust the number of active workers based on real-time traffic load. This ensures that the system maintains performance during peak request periods while minimizing costs during low-traffic periods. This balance of performance and efficiency is a critical aspect of managing production ML infrastructure in an enterprise Databricks environment.

Exam trap

Candidates often assume auto_scale refers to cluster resizing or data processing parallelism, failing to realize it specifically governs the number of compute instances serving the model endpoint based on traffic.

4
MCQhard

A fraud-detection model is registered in Unity Catalog and deployed to a Model Serving endpoint. Compliance requires that every prediction be traceable to the exact model version and the request that produced it. Which combination of Databricks features should you configure to meet this requirement?

A.Enable inference table logging on the endpoint and register the model version in Unity Catalog with a descriptive comment.
B.Attach a model signature to the registered model and enable request logging in the serving client application.
C.Configure the endpoint with a scale-to-zero policy and log the endpoint configuration to a Delta table daily.
D.Enable MLflow autologging in the training notebook and set the model version stage to Production.
AnswerA

Inference tables record each request, response, timestamp, and the served model version, while Unity Catalog provides the lineage and version identity of the registered model. Together they let an auditor trace any prediction back to the exact model version and the originating request, which is precisely the compliance requirement.

Why this answer

Traceability of individual predictions requires server-side capture of the request, response, and serving model version, which is what inference tables provide, combined with Unity Catalog's governed model version identity. Training-time logging, endpoint configuration snapshots, and client-side logs each capture only part of the picture and cannot reconstruct the full per-prediction audit trail.

Exam trap

The trap here is assuming that MLflow autologging or a model signature provides runtime prediction traceability, when those features operate at training time or schema validation only.

5
Multi-Selectmedium

A team is preparing to promote a new model version to production in the MLflow Model Registry. They must ensure the model can be served with a consistent environment across staging and production and that dependency drift is detected before promotion. Which TWO practices should they follow? (Choose two.)

Select 2 answers
A.Store only the model's pickle file and reconstruct dependencies manually on each environment, since pickle serialization is environment-independent.
B.Log the model with its conda environment and requirements files so the exact dependency versions are captured as part of the model version.
C.Rely on the serving endpoint to install the latest available versions of each library at deployment time so the environment stays current.
D.Verify the logged environment files against the target serving environment before promotion, resolving any version mismatches in the model's dependency specification.
E.Pin the serving cluster's runtime to a different Databricks Runtime version than training used, to validate cross-version compatibility during staging.
AnswersB, D

Logging the conda environment and requirements files records the exact library versions used at training time as part of the model artifact. When the same model version is served in staging and production, the environment can be reconstructed from these files, which is the foundation for consistent, reproducible serving and for detecting any drift from the recorded versions.

Why this answer

Capturing the exact dependency environment with the model version and then validating that environment against the target serving environment before promotion together ensure consistent serving and catch drift early. Floating to latest versions, mismatching runtimes, or relying on bare pickles all break reproducibility or hide drift.

Exam trap

The trap here is thinking that a serialized model file alone guarantees reproducible serving, when the surrounding dependency environment must also be captured and verified.

6
MCQmedium

Your team is experiencing 'training-serving skew' where model performance in production is significantly lower than during training. Which approach should you prioritize to mitigate this issue?

A.Increase the amount of historical training data used in the model.
B.Use the Databricks Feature Store to unify feature engineering pipelines.
C.Switch the model architecture to a deeper neural network.
D.Retrain the model every hour to keep it fresh with recent data.
AnswerB

The Feature Store enforces consistent data processing by design. By using the same feature definitions for both training and serving, you ensure that the input features are computed identically in both environments, which is the industry-standard method for resolving training-serving skew in production machine learning.

Why this answer

Training-serving skew often occurs because the data transformation logic used during model training differs from the logic applied during real-time inference. By utilizing the Databricks Feature Store, you can centralize your feature engineering code. This ensures that the exact same transformation code and look-up logic are utilized for both offline training and online inference, effectively eliminating the discrepancy between the two environments.

Exam trap

Candidates often look for retraining frequency adjustments or hyperparameter tuning, failing to recognize that training-serving skew is primarily caused by divergent feature pipelines.

7
MCQmedium

Which component of MLflow is responsible for keeping track of the different versions of a model as it moves from development to testing and production?

A.MLflow Tracking
B.MLflow Projects
C.MLflow Model Registry
D.MLflow Recipes
AnswerC

The Model Registry provides a centralized store for managing the full lifecycle of MLflow models. It supports versioning, stage transitions, and annotations. It is the primary tool in Databricks for managing production deployments, ensuring that the correct model versions are used in the correct environments at all times.

Why this answer

The MLflow Model Registry is the central repository for model versioning. It allows teams to manage the lifecycle of a model by transitioning versions through different stages. This is a critical MLOps function because it ensures that production environments are always pointing to a known, stable version of the model, while allowing developers to continue iterating on new versions without disrupting the live service.

Exam trap

Candidates often confuse the MLflow Tracking Server with the Model Registry, failing to distinguish between experiment logging and the centralized management of model production versions.

8
Multi-Selecthard

Your team uses Databricks for machine learning and needs to ensure that model training is fully automated and reproducible. Which THREE of the following are necessary components for a production-grade automated ML pipeline?

Select 3 answers
A.Git integration to version control the training notebooks and scripts.
B.Databricks Workflows to schedule and orchestrate pipeline steps.
C.MLflow Model Registry to manage model versions and deployment lifecycle.
D.Manual approval steps for every single model training run.
E.Using the same cluster for both development and production tasks to minimize complexity.
AnswersA, B, C

Version control is fundamental to MLOps. It allows teams to track changes, collaborate, and revert to known good states. Without Git, it is impossible to audit the evolution of training logic or ensure that the code running in production is exactly the same as what was tested during the CI process.

Why this answer

An automated pipeline requires a workflow orchestrator (Workflows), a centralized registry (MLflow), and version-controlled logic (Git). These three components work together to provide a seamless transition from code development to production deployment. This triad ensures that every step is reproducible, monitored, and audit-ready, which is the foundational requirement for any mature MLOps practice operating at enterprise scale.

Exam trap

Candidates often include 'manual monitoring' or 'local testing' as key components, failing to identify the three pillars of automation: Git, Workflows, and the Model Registry.

9
MCQhard

You are monitoring a model served on Databricks Model Serving. You need to detect data drift in the incoming requests without delaying predictions. Which approach should you use?

A.Add a pre-processing step in the model's predict function that computes drift metrics on each request.
B.Configure the endpoint to call an external monitoring service synchronously before returning predictions.
C.Use Databricks SQL dashboards to query the model's training data and compare it to live predictions in real time.
D.Enable inference logging on the endpoint and analyze the logged requests asynchronously using a Databricks job that computes drift metrics.
AnswerD

Inference logging captures the request payloads and predictions to a Delta table. You can then run an asynchronous job to compute drift metrics, such as population stability index or KL divergence, comparing recent data to a baseline. This does not add latency to predictions because logging is decoupled from the serving path, making it the recommended approach.

Why this answer

Inference logging on Model Serving captures request and response data to a Delta table without adding latency to the prediction path. An asynchronous job can then analyze this logged data to compute drift metrics, enabling detection without impacting performance. Synchronous or in-model computations are impractical due to latency and single-request limitations.

Exam trap

The trap here is assuming that drift detection must happen inline with predictions, but it should be decoupled to avoid latency and allow windowed analysis.

10
MCQmedium

Refer to the exhibit. You are managing the 'revenue_forecast' model in the registry. A colleague wants to deploy this version to production. Which Databricks command or process is required to move version 4 to the 'Production' stage while ensuring existing production models remain unaffected?

A.Delete the existing Production version and then promote version 4 to Production.
B.Use the MLflow transition_model_version_stage API with archive_existing_versions=True.
C.Directly update the 'stage' tag in the model JSON definition to 'Production'.
D.Reregister the model as a new model name to avoid conflicts with version 3.
AnswerB

This method is the programmatic way to promote a new version while safely archiving the previous one. Setting the parameter to True ensures that the registry remains clean and that there is no ambiguity regarding which model version is currently serving production traffic for the 'revenue_forecast' application.

Why this answer

In Databricks, using the 'transition_model_version_stage' method with 'archive_existing_versions=True' is the standard practice for seamless deployments. This automatically moves the previous production version to 'Archived', maintaining a clear audit trail. Proper lifecycle management ensures that only one version is active in production at a time, preventing conflicts and ensuring that the most current, verified model is the one serving live traffic.

11
MCQeasy

A data scientist has trained a scikit-learn model and logged it with MLflow. They now want to register this model in the MLflow Model Registry and transition it to 'Production' to be served via Databricks Model Serving. Which of the following is a prerequisite for registering the model?

A.The model must be logged using the MLflow Tracking API and have a valid run ID.
B.The model must be logged to an MLflow experiment that is associated with a registered model.
C.The model must be logged with a signature.
D.The model must be approved by a workspace admin.
AnswerA

To register a model, you need a model artifact that is associated with a run in MLflow Tracking. The run ID provides the lineage and allows the registry to reference the model's source. Without a valid run, registration cannot proceed.

Why this answer

Registering a model in MLflow Model Registry requires that the model artifact is linked to an MLflow run. The run ID establishes the model's origin and enables versioning. Other options are either optional or incorrect; signatures, experiment associations, and admin approvals are not required for registration.

Exam trap

The trap here is assuming that a model signature or admin approval is needed to register a model, when actually only a valid MLflow run is required.

12
MCQeasy

You are monitoring a model deployed to a Databricks Model Serving endpoint. You need to track the distribution of input features and model predictions over time to detect data drift. Which Databricks feature should you use?

A.Inference tables
B.Delta Live Tables
C.MLflow tracking server logs
D.Databricks SQL dashboards
AnswerA

Inference tables automatically capture the input features and model predictions for each request to a Model Serving endpoint. They store the data in a Delta table, enabling you to analyze feature distributions and detect drift over time. This is the built-in solution for monitoring model serving traffic on Databricks.

Why this answer

Inference tables automatically log the input features and predictions for each request to a Model Serving endpoint, storing them in a Delta table. This allows you to monitor feature distributions and detect data drift over time using SQL or dashboards, making it the correct choice for this scenario.

Exam trap

The trap here is confusing training-time tracking (MLflow) with inference-time monitoring; only inference tables capture production request data for drift detection.

13
MCQmedium

Refer to the exhibit. An MLOps engineer is reviewing a JSON object representing a model in the Databricks Model Registry. The engineer wants to promote this model to the 'Production' stage using the MLflow Python API. Which command is correct?

A.client.transition_model_version_stage(name='revenue_forecast', version=5, stage='Production')
B.client.update_model_version(name='revenue_forecast', version=5, new_stage='Production')
C.mlflow.register_model(model_uri='revenue_forecast/5', stage='Production')
D.client.set_model_stage(model='revenue_forecast', ver=5, to='Production')
AnswerA

This method is the correct MLflow API call for changing the stage of a registered model. By specifying the model name, version, and target stage, the engineer programmatically moves the model through the lifecycle, which is a fundamental requirement for automated deployment pipelines in Databricks.

Why this answer

To transition a model version in MLflow, the `transition_model_version_stage` function is the standard method. It requires the model name, the specific version, and the target stage string. This API call is critical for programmatic CI/CD pipelines, allowing engineers to automate the promotion process based on validation results, thereby reducing manual effort and ensuring consistent deployment procedures across the organization.

Exam trap

Candidates often guess the API method name, confusing it with generic MLflow logging methods or incorrectly assuming they need to delete and re-register the model to change its stage.

14
MCQhard

When designing a robust MLOps pipeline for high-stakes financial applications, why should you prioritize 'reproducibility' over 'speed' during the deployment phase?

A.Speed is irrelevant in all machine learning contexts.
B.Reproducibility is required for compliance, auditability, and reliable debugging.
C.Databricks does not support fast deployment pipelines.
D.Speed increases the risk of data leakage during training.
AnswerB

Financial regulators require evidence of how model decisions are made. If you cannot reproduce the training process, you cannot guarantee the integrity of the model. Furthermore, debugging production issues is impossible without being able to recreate the exact environment, artifacts, and data state that produced the problematic output.

Why this answer

Reproducibility ensures that a model can be recreated exactly using the same code, environment, and data. In high-stakes fields like finance, being able to audit every decision is a regulatory requirement. While speed is useful, a fast deployment of an un-reproducible model creates significant business risk.

If a model fails, you must be able to recreate the state to identify and fix the bug to remain compliant.

Exam trap

Candidates often choose speed over reproducibility because agile deployment sounds modern, overlooking strict regulatory and auditing requirements inherent in high-stakes financial environments.

15
MCQmedium

A data science team at a retail company has registered a demand forecasting model in the MLflow Model Registry. The model version is currently in the 'Staging' stage and has been validated by the QA team. Before promoting it to 'Production', the ML engineer wants to ensure that the model's performance does not degrade when serving live traffic. They decide to deploy the model to a small percentage of production traffic while continuing to serve the existing model. Which Databricks feature should they use to achieve this?

A.Databricks Model Serving with traffic splitting
B.Databricks Jobs with conditional task execution
C.MLflow Model Registry webhooks
D.MLflow Projects with Docker environments
AnswerA

Databricks Model Serving supports traffic splitting, allowing you to route a percentage of requests to a new model version while the rest go to the existing version. This enables safe canary deployments and A/B testing. By configuring traffic splitting on the endpoint, the team can gradually shift traffic and monitor performance without impacting all users.

Why this answer

Databricks Model Serving natively supports traffic splitting, which lets you distribute incoming requests across multiple model versions on the same endpoint. This is ideal for canary deployments: you can send a small fraction of traffic to the new version, monitor metrics, and then increase the percentage. Other options lack the real-time routing capability required to test a model under live production load without affecting all users.

Exam trap

The trap here is confusing model deployment automation (webhooks, jobs) with traffic management, which requires a serving feature that can split requests between versions.

16
MCQmedium

You maintain a Databricks ML pipeline that trains a model nightly and registers new versions in MLflow Model Registry. A downstream batch scoring job in another workspace loads the model by stage. Auditors require that every production scoring run can be traced back to the exact training data snapshot and code commit. Which approach best satisfies this requirement?

A.Use mlflow.log_input with a Delta table dataset, log the Git commit as a run tag, and register the model version from that run so the run ID links model, data, and code.
B.Store the training data path in the model's signature and use the workspace notebook revision to identify the code commit.
C.Set the model version's description to the Git commit hash and rely on the registered model name to identify the data snapshot.
D.Log the training run with mlflow.log_param for the Git commit and log the Delta table version as a tag, then register the model version from that run.
AnswerA

mlflow.log_input records a dataset entity with its source and digest on the run, and the run ID becomes the immutable link between the model version, the exact data snapshot, and the logged Git commit tag. Because MLflow automatically stores the run ID on the model version, auditors can traverse from a scored model back to the precise training inputs and code revision without relying on human-maintained metadata.

Why this answer

The requirement is an immutable, queryable chain from a production model version to the exact training data snapshot and code revision. Logging the dataset via mlflow.log_input and the Git commit as a run tag, then registering from that run, creates that chain because the model version stores the run ID, and the run stores the dataset digest and code tag. Manual descriptions, parameters, signatures, or notebook revisions are mutable or incomplete and cannot satisfy an audit.

Exam trap

The trap here is assuming that any logged metadata such as a description or parameter creates an auditable link, when only the run ID and logged dataset entities provide an immutable, queryable chain.

17
MCQmedium

You are using MLflow to track experiments for a computer vision model. After several runs, you notice that the training loss is not being logged correctly because the metric name contains a space. What is the recommended way to log metrics with names that include spaces?

A.Use the `mlflow.log_metric` function with the space included; MLflow will automatically sanitize it.
B.Replace spaces with underscores or hyphens before logging.
C.Log the metric as a tag instead of a metric.
D.Use a custom MLflow plugin to allow spaces in metric names.
AnswerB

MLflow metric and parameter names must be valid identifiers; spaces are not allowed. The recommended practice is to sanitize names by replacing spaces with underscores or hyphens. This ensures the metrics are logged correctly and can be queried and visualized without errors in the MLflow UI and API.

Why this answer

MLflow enforces strict naming rules for metrics and parameters, disallowing spaces. The correct approach is to replace spaces with underscores or hyphens before logging. This maintains compatibility with the MLflow tracking server, UI, and API, and avoids logging errors.

Other options either do not solve the problem or introduce unnecessary complexity.

Exam trap

The trap here is thinking that MLflow will automatically handle spaces in metric names or that tags are a suitable substitute for numeric metrics.

18
MCQmedium

You are preparing a feature-engineering job that reads from a Delta table and writes a feature table. The job must run on a schedule and be idempotent so that reruns after a failure do not duplicate data. Which approach should you use?

A.Use a Databricks Job with a notebook that performs a MERGE INTO the feature table using a deterministic key and a timestamp condition.
B.Use a Databricks Job with a notebook that uses DataFrame.write.mode("overwrite") to replace the entire feature table on each run.
C.Use a Databricks Job with a Delta Live Tables pipeline that writes to a streaming table with a defined primary key and applies AUTO CDC.
D.Use a Databricks Job with a notebook that appends new rows and relies on a downstream deduplication step in the serving layer.
AnswerA

MERGE INTO with a deterministic key and a condition ensures that on rerun, existing rows are updated rather than inserted again, making the operation idempotent. This is the standard pattern for scheduled jobs that must tolerate retries and failures without creating duplicates, and it works with Delta Lake's ACID guarantees.

Why this answer

A MERGE INTO operation keyed on a deterministic identifier and a condition ensures that repeated runs update existing records instead of inserting duplicates, satisfying idempotency. This is the recommended pattern for scheduled feature-engineering jobs in Databricks that must be resilient to retries. The other options either rely on downstream fixes, destroy incremental state, or use mechanisms not designed for idempotent batch writes.

Exam trap

The trap here is assuming that any Delta write is automatically idempotent, when in fact append operations will duplicate rows on rerun unless a MERGE or similar upsert logic is used.

19
MCQhard

A machine learning engineer is using MLflow Tracking to log metrics and artifacts for a deep learning model. They notice that the training run logs a large number of metrics (e.g., loss per batch) and want to reduce the storage footprint and improve query performance. They also need to retain the ability to compare runs and reproduce results. Which of the following actions is most appropriate?

A.Disable metric logging entirely and rely on TensorBoard for visualization.
B.Log only the final epoch metrics and use MLflow's step parameter to log intermediate metrics less frequently.
C.Use a different tracking server with a NoSQL backend to handle the high volume of metrics.
D.Store metrics in a separate file and log it as an artifact instead of using mlflow.log_metric.
AnswerB

Logging metrics with the step parameter allows you to record a time series, but logging every batch can create a huge number of metric entries. By logging only final epoch metrics or reducing frequency, you decrease storage and speed up queries. MLflow still stores the history if needed, but the volume is manageable. This balances reproducibility with efficiency.

Why this answer

MLflow's log_metric function can accept a step parameter to record metrics over time, but logging every batch creates a large number of entries. Reducing the logging frequency or only logging final epoch metrics significantly decreases storage and improves query speed while still providing enough data for comparison and reproducibility. Other options either lose MLflow's tracking capabilities or introduce unsupported configurations.

Exam trap

The trap here is assuming that any reduction in logged metrics requires abandoning MLflow tracking, when in fact you can simply log less frequently using the step parameter.

20
MCQmedium

A data scientist has completed hyperparameter tuning with Hyperopt on Databricks and now needs to register the best-performing model to the MLflow Model Registry, including its signature and input example, so that downstream scoring jobs can validate incoming data. The training script uses MLflow autologging. Which approach most reliably captures the signature and input example during registration?

A.Use the MLflow Client's update_model_version() method after registration to attach the signature and input example as tags on the model version.
B.Call mlflow.register_model() directly on the run URI returned by the tuning run, because autologging automatically infers and stores the signature and input example.
C.Explicitly call mlflow.log_model() (or the flavor-specific log_model) with the signature and input_example arguments inside the training function, then register that logged model.
D.Register the model from the MLflow experiment's artifact location by copying the MLmodel file and manually editing it to include the signature and input example fields.
AnswerC

Passing signature and input_example to a flavor-specific log_model call records both artifacts with the model version. This is the deterministic way to guarantee the signature schema and a representative input example are attached, which downstream scoring jobs can then use to validate incoming payloads before invoking the model.

Why this answer

The signature and input example are artifacts attached at model logging time, so the reliable path is to pass them explicitly to the flavor-specific log_model call used by the training function. Registering afterward from the run URI, editing tags, or hand-editing the MLmodel file does not produce the complete, tooling-readable artifacts needed for downstream data validation.

Exam trap

The trap here is assuming that MLflow autologging always captures both a signature and an input example, when in practice the input example must be supplied explicitly.

21
MCQeasy

A data scientist has trained a model and logged it with MLflow. They now need to deploy it as a real-time endpoint on Databricks Model Serving. The model requires a custom Python library that is not available in the default environment. What is the correct way to ensure the library is available when serving the model?

A.Install the library on the driver node of the cluster used for serving, and restart the cluster.
B.Use a custom container image for the model endpoint and install the library at runtime via a startup script.
C.Add the library to the Databricks workspace's global init script so it is available to all clusters.
D.Include the library in the model's conda environment or requirements file when logging the model with MLflow.
AnswerD

When logging a model with MLflow, you can specify a conda environment or a requirements file that lists all dependencies. Databricks Model Serving uses this information to create the serving environment, ensuring the custom library is installed. This is the standard and supported method for managing model dependencies in serving.

Why this answer

The correct way is to include the custom library in the model's conda environment or requirements file when logging with MLflow. Databricks Model Serving uses this specification to build the serving environment, ensuring the library is available. Other methods like cluster-level installations or runtime scripts do not apply to serverless serving endpoints and can introduce reliability issues.

Exam trap

The trap here is assuming that Model Serving endpoints run on clusters you can configure, but they are serverless and dependencies must be declared in the model artifact.

22
MCQmedium

Refer to the exhibit. If the 'train_model' task fails, what happens to the 'evaluate_model' task in this Databricks Job?

A.It executes regardless of the failure.
B.It is skipped because the dependency is not met.
C.It attempts to run using the last successful model artifact.
D.It retries the 'train_model' task automatically until success.
AnswerB

The 'depends_on' field creates a strict prerequisite. When the parent task fails, the DAG execution stops or skips the child task by design. This mechanism protects the pipeline from executing logic based on incomplete or invalid model artifacts, maintaining the reliability of the overall MLOps training workflow.

Why this answer

In Databricks Jobs, a task that depends on another will only execute if the parent task finishes successfully. If 'train_model' fails, the dependency condition for 'evaluate_model' is not met, causing the job to stop or skip the downstream task. This behavior is intentional to prevent evaluating a non-existent or corrupted model, ensuring pipeline integrity and avoiding wasted computational resources on dependent tasks.

Exam trap

Candidates assume the job continues regardless of failure, not realizing that Databricks Jobs default to a strict dependency model where downstream tasks are aborted if the parent fails.

23
MCQmedium

A data science team is transitioning from manual model training to automated pipelines in Databricks. They require a mechanism to track model lineage, versions, and stage transitions programmatically. Which Databricks component best satisfies this requirement?

A.Databricks Jobs Scheduler
B.MLflow Model Registry
C.Databricks Repos
D.Delta Lake
AnswerB

The Model Registry is specifically engineered to handle model versioning, stage transitions (e.g., Staging to Production), and lineage tracking. It provides the necessary APIs to automate these workflows within CI/CD pipelines, ensuring that model deployment follows a robust and governed process in Databricks environments.

Why this answer

MLflow Model Registry provides a centralized model store, set of APIs, and UI for managing the full lifecycle of MLflow models. It enables versioning, model stage transitions, and lineage tracking, which are critical for MLOps maturity. By using the Model Registry, teams ensure reproducibility and governance, allowing for seamless promotion of models from staging to production while maintaining an audit trail of changes and deployment history.

Exam trap

Candidates often confuse MLflow Tracking with MLflow Model Registry. They select Tracking because it records metrics, but it lacks the lifecycle management and stage transition capabilities required for formal model promotion.

24
MCQeasy

When implementing a Feature Store in Databricks, what is the primary benefit of using a Feature Table compared to a standard Delta table for feature engineering?

A.Feature Tables automatically convert all data into Parquet format.
B.They enable automated point-in-time joins for training data generation.
C.They provide faster query performance than standard Delta tables.
D.They allow users to write SQL queries directly to the feature data.
AnswerB

Point-in-time joins are critical for avoiding training-serving skew and data leakage. Feature Tables store metadata that allows the Feature Store to perform these joins automatically, ensuring that the features used for training are temporally consistent and truly represent the state at the time of observation.

Why this answer

Feature Tables provide metadata tracking and lineage, linking features directly to the models that consume them. This metadata enables automated point-in-time joins, preventing data leakage during training by ensuring features are sampled as they existed at a specific timestamp. This level of automation is not available with standard Delta tables, making Feature Tables essential for building reproducible, production-ready machine learning pipelines.

Exam trap

Candidates mistakenly believe that Feature Tables are just a storage format for performance, ignoring the critical architectural capability of point-in-time joins which prevents data leakage during training.

25
MCQeasy

A data scientist has developed a scikit-learn model and wants to deploy it as a REST API endpoint on Databricks Model Serving. They have logged the model with MLflow and registered it in the MLflow Model Registry. The model requires a specific Python library that is not pre-installed in the Databricks Model Serving environment. What should the data scientist do to ensure the library is available when the model is served?

A.Package the library as a wheel file and upload it to DBFS, then reference it in the serving endpoint configuration.
B.Add the library to the model's conda environment and log the model with that environment.
C.Modify the cluster's Spark configuration to include the library path.
D.Install the library on all cluster nodes using an init script.
AnswerB

Databricks Model Serving uses the conda environment specified in the MLflow model to install dependencies. By including the required library in the conda environment and logging the model with it, the serving environment will install the library automatically. This ensures the model has all necessary dependencies at runtime.

Why this answer

Databricks Model Serving builds the serving environment based on the conda environment recorded with the MLflow model. To include a custom library, you must add it to that conda environment before logging the model. When the model is deployed, the serving infrastructure will create an environment with the specified dependencies.

Other methods like init scripts or Spark config do not apply to the serverless serving environment.

Exam trap

The trap here is assuming that Model Serving uses clusters or can reference external files for dependencies, when it actually relies on the conda environment packaged with the model.

26
MCQmedium

Your company uses MLflow tracking. You want to query all experiments that achieved a specific accuracy threshold across multiple teams. What is the most efficient way to achieve this?

A.Export all experiment metadata to a CSV and use Excel to filter for the desired accuracy.
B.Use the 'mlflow.search_runs()' API to programmatically filter runs based on the accuracy metric.
C.Manually browse through each experiment folder in the MLflow UI to find the metrics.
D.Query the underlying Delta tables in the MLflow experiment storage location directly.
AnswerB

The search_runs API is the standard and most efficient way to query MLflow experiment data. It supports filtering by metrics, parameters, and tags, enabling users to quickly retrieve relevant information programmatically. This method integrates perfectly with automated analysis tools, CI/CD pipelines, and centralized reporting dashboards.

Why this answer

The MLflow Search API (search_runs) is specifically designed to query experiments based on metrics, parameters, and tags. By using the Python MLflow client to filter runs, you can programmatically extract insights across different workspaces. This avoids manual searching, allows for automated report generation, and facilitates governance by allowing you to easily identify top-performing models or flag models that don't meet corporate performance standards.

Exam trap

Candidates often suggest using the UI to manually filter runs, which is inefficient for large-scale enterprise environments compared to the programmatic 'search_runs' API.

27
MCQmedium

You are migrating a legacy ML pipeline to Databricks. You need to ensure that the feature engineering logic used during training is identical to the logic used during real-time inference. What is the recommended approach?

A.Copy the feature transformation code into both the training notebook and the inference service.
B.Use the Databricks Feature Store to define and log features.
C.Store the features in a Parquet file and load it during both training and inference.
D.Implement the transformation logic as a SQL view in the Hive Metastore.
AnswerB

The Feature Store enables the packaging of features with their associated transformation logic. When you use the Feature Store for training, it automatically logs the transformations. During inference, you can then retrieve the features using the same ID, ensuring that the exact same logic is applied to the input data.

Why this answer

Consistency between training and inference (the 'training-serving skew') is a primary cause of model failure. The Databricks Feature Store acts as the single source of truth for feature definitions. By using it, you ensure that the same transformations are applied during both training and inference, eliminating discrepancies that arise from re-implementing logic in different languages or frameworks, which is critical for model reliability.

28
MCQmedium

You are designing a strategy for monitoring model performance after deployment. Which of the following is the most important indicator that a model requires retraining?

A.An increase in the number of concurrent inference requests.
B.A significant drop in model prediction accuracy on production data.
C.The expiry of the model's 'Production' stage status.
D.A minor change in the underlying Databricks runtime version.
AnswerB

Model accuracy is the bottom-line metric for performance. A drop in accuracy indicates that the relationship between inputs and outputs has changed, implying that the model is no longer reflecting the current reality of the data. This is the clearest, most urgent signal that the model requires retraining.

Why this answer

Performance degradation, often caused by model drift or data drift, is the primary driver for retraining. While monitoring system metrics like latency is important, the predictive quality of the model is the ultimate metric. Detecting a significant drop in accuracy or precision indicates that the model is no longer meeting business requirements, triggering the need for a new training cycle in an automated MLOps workflow.

29
MCQmedium

You are designing a model retraining strategy. What is the most reliable way to trigger a retraining job based on model performance degradation?

A.Schedule the retraining job to run every Monday regardless of current model performance.
B.Use Databricks Workflows to trigger retraining based on drift detection alerts.
C.Send an email alert to the lead data scientist to manually kick off the job.
D.Write a cron job in the notebook that calculates performance every minute.
AnswerB

This event-driven approach ensures that retraining only occurs when it is objectively needed. By triggering Workflows based on drift detection, you maintain the model's relevance to the current data distribution without wasting compute budget on redundant retraining jobs when the model is still performing adequately.

Why this answer

Linking monitoring metrics to workflow orchestration is the standard MLOps pattern for automated retraining. By configuring a monitor that triggers a Databricks Job when performance drifts below a defined metric (e.g., accuracy), you create a closed-loop system. This ensures that models are updated only when necessary, minimizing cost while maintaining high model performance and reducing manual intervention in the model lifecycle.

Exam trap

Test-takers often assume models retrain automatically based solely on elapsed time, overlooking that event-driven workflows triggered by drift alerts are the best practice.

30
Multi-Selecthard

Which TWO actions are required to properly implement MLflow Model Registry stages and governance for a machine learning project?

Select 2 answers
A.Assign the 'Can Manage' permission on the Model Registry to all users.
B.Transition model versions between stages like 'Staging' and 'Production'.
C.Apply ACLs to specific models to restrict access and promote actions.
D.Keep all model versions in the 'None' stage for maximum flexibility.
E.Hard-code the model version string into the production application code.
AnswersB, C

Using stages helps manage the lifecycle of a model effectively. It provides a clear signal to downstream applications about which model versions are approved for specific environments, facilitating a clean handoff from development to production and allowing for programmatic automated deployment pipelines.

Why this answer

Implementing Model Registry stages ensures that only validated models move from development to production. By controlling access permissions, organizations enforce a separation of duties, ensuring that data scientists can register models, while only authorized engineers can promote them. This gated workflow is foundational for regulatory compliance and auditability in enterprise MLOps, preventing unauthorized model deployments into downstream production environments.

Exam trap

Candidates frequently select 'registering models' as the governance action, but registration is a development task. Governance specifically requires stage transitions and ACLs to enforce separation of duties.

31
MCQhard

Refer to the exhibit. Which security configuration is the likely culprit for this job failure?

A.Missing cluster-level library permissions.
B.Insufficient Unity Catalog permissions for the job identity.
C.The table is stored in an encrypted S3 bucket.
D.The Delta table has reached its version limit.
AnswerB

The job is executing under a specific identity (service principal). If that identity lacks the necessary 'SELECT' or 'USE' grants in Unity Catalog for the specified schema, access to the underlying Delta table will be blocked, resulting in the reported permission error during the job execution.

Why this answer

This error points to a failure in Unity Catalog access control. In Databricks, when a job attempts to read from a table, it requires specific 'SELECT' permissions on the table and 'USE' permissions on the catalog and schema. If the job service principal is not granted these roles, the query will fail regardless of the workspace-level permissions assigned to the user.

Exam trap

Students frequently blame cluster node sizing or notebook syntax errors, missing the governance layer where job service principals lack Unity Catalog permissions.

32
MCQmedium

A fraud detection model was deployed to a Databricks Model Serving endpoint. Over the past week, the endpoint's p99 latency has increased from 45 ms to 320 ms, but the model's predictions remain accurate. The serving logs show that each request now includes a larger JSON payload with additional transaction metadata. What is the most likely cause of the degradation?

A.The serving endpoint is hitting the maximum number of concurrent requests and queuing requests, which adds wait time.
B.The model's feature vector has grown, and the scoring function is performing expensive feature transformations on the raw payload before inference.
C.The model artifact is being re-downloaded from the MLflow Model Registry on every request due to a misconfigured artifact cache.
D.The model's underlying compute cluster has been scaled down, reducing available CPU resources for inference.
AnswerB

When the request payload includes extra fields, the model's predict function or preprocessing logic may compute additional features, increasing CPU time per request. This directly explains the latency increase while accuracy stays stable. The other options either would affect accuracy or are not supported by the symptom of larger payloads.

Why this answer

The increased payload size likely triggers additional feature engineering or parsing inside the model's scoring logic, raising per-request CPU time. Since accuracy is unchanged, the model itself is fine; the bottleneck is the preprocessing step. The other explanations either contradict the payload-size correlation or would impact accuracy or availability.

Exam trap

The trap here is assuming that latency degradation always indicates infrastructure scaling problems, ignoring that request payload changes can increase per-request processing time.

33
MCQeasy

When deploying a model as a real-time REST API endpoint on Databricks, which service should be used to manage the serving infrastructure, scaling, and availability?

A.Databricks SQL Warehouse
B.Databricks Model Serving
C.Apache Spark Streaming
D.Databricks File System (DBFS)
AnswerB

Model Serving is the dedicated Databricks product for deploying models as REST APIs. It handles the complexities of scaling, security, and availability, allowing teams to focus on model performance rather than infrastructure management. This is the industry-standard way to expose models within the Databricks ecosystem.

Why this answer

Databricks Model Serving is the managed service designed specifically to deploy models as low-latency, high-availability REST endpoints. It abstracts the underlying infrastructure, providing automatic scaling, version management, and health monitoring. This service is essential for production workloads where reliability and ease of maintenance are paramount, as it removes the burden of manual server administration and infrastructure configuration from the data science team.

Exam trap

Candidates confuse Databricks Model Serving with standard batch jobs or cluster deployments, failing to choose the managed real-time endpoint service.

34
MCQmedium

When deploying a model to a Databricks Model Serving endpoint, how can you ensure the model scales automatically based on traffic demand?

A.Manually adjust the number of instances in the endpoint configuration every time traffic increases.
B.Set the endpoint to a fixed number of instances that is high enough to handle peak traffic.
C.Configure auto-scaling settings in the Model Serving endpoint definition to handle traffic fluctuations.
D.Use a load balancer to redirect traffic to different model endpoints depending on the time of day.
AnswerC

Auto-scaling is a core feature of Databricks Model Serving. By enabling this configuration, the platform automatically manages the allocation of compute resources to match incoming traffic patterns. This provides a balance between high availability during peaks and cost optimization during periods of lower utilization without manual intervention.

Why this answer

Databricks Model Serving utilizes auto-scaling compute resources to handle fluctuations in traffic. By configuring the serving endpoint with appropriate auto-scaling parameters, the platform dynamically adjusts the number of concurrent instances based on request load. This ensures that the application maintains low latency under heavy load while remaining cost-effective during periods of low activity, which is essential for managing production-level model inference at scale.

Exam trap

Candidates often confuse manual endpoint configuration with auto-scaling, mistakenly believing they must manually add nodes or adjust cluster sizes during traffic spikes in the Databricks UI.

35
MCQhard

Refer to the exhibit. Why is the missing model signature considered an MLOps risk in a production environment?

A.It prevents the model from being saved to the local file system.
B.It restricts the model to only be deployed on CPU clusters.
C.It makes it difficult to perform automated input validation at inference time.
D.It slows down the training process by adding extra validation overhead.
AnswerC

Without a signature, the serving system doesn't know the expected input types or shapes. This prevents automated validation, where the system checks data against the schema before passing it to the model. This increases the risk that malformed data will cause the model to crash or produce garbage results.

Why this answer

Model signatures define the expected input schema. Without them, the serving engine cannot validate that incoming data matches what the model expects, which leads to runtime errors or silent, incorrect predictions. In production, this lack of validation makes it impossible to distinguish between data errors and model logic errors, creating significant operational risks and slowing down root cause analysis during incidents.

Exam trap

Candidates often think missing signatures only affect documentation or metadata, failing to realize they are critical for the serving infrastructure to perform automated input validation.

36
MCQmedium

A data scientist has registered a new model version in MLflow Model Registry and wants to validate it against a golden dataset before promoting it. The validation notebook must run automatically whenever a new version is registered in the 'ChurnModel' registry model, and the result should block promotion if the validation fails. Which approach should the data scientist use?

A.Configure a webhook on the MLflow Model Registry that triggers a Databricks job running the validation notebook, and use the job's success/failure to gate promotion.
B.Create a Databricks job with a file arrival trigger that monitors the MLflow artifact store for new model files.
C.Enable automatic model promotion in the registry so that any newly registered version is immediately moved to Production after passing built-in validation.
D.Use MLflow Projects to package the validation notebook and run it manually each time a new version is registered.
AnswerA

MLflow Model Registry webhooks fire on registry events such as MODEL_VERSION_CREATED, allowing an external service (here, a Databricks job) to run validation automatically. The job's outcome can be used to gate promotion, satisfying the requirement for automatic validation upon registration.

Why this answer

MLflow Model Registry webhooks are designed to notify external systems when registry events occur, such as a new model version being created. By configuring a webhook that calls a Databricks job, the validation notebook runs automatically. The job's success or failure can then be used to decide whether to promote the version, providing the required gating mechanism.

Exam trap

The trap here is assuming MLflow Model Registry has built-in automatic validation or promotion, when in fact it relies on webhooks and external orchestration for such workflows.

37
MCQhard

A nightly Databricks job trains a model, registers a new version in Unity Catalog, and then updates a Model Serving endpoint that serves the 'champion' alias. The endpoint must switch to the new version only after the job's validation step passes. Which approach correctly enforces this?

A.Configure the endpoint with a scale-out policy that adds replicas whenever a new model version is registered.
B.Delete the previous model version after training so the endpoint is forced to pick up the newest registered version.
C.Have the job update the endpoint configuration to reference the new model version number directly after training completes.
D.Assign the 'champion' alias to the new version only inside the validation-passed branch of the job, and leave the endpoint configured to serve the alias.
AnswerD

Because the endpoint serves the alias, moving the alias to the validated version is the promotion step, and doing it only after validation passes guarantees the endpoint never serves an unvalidated version. Rollback is a single alias reassignment back to the prior version.

Why this answer

Decoupling promotion from deployment is the key: the endpoint always serves the 'champion' alias, and the job reassigns that alias only after validation succeeds. This makes the alias the single source of truth for what is in production and keeps rollback to a one-step alias change, while the validation branch ensures unvalidated versions never reach the endpoint.

Exam trap

The trap here is thinking that registering a new model version automatically causes a serving endpoint to switch to it, when endpoints only change when their referenced alias or version changes.

38
MCQhard

A team has a production Databricks Model Serving endpoint for a churn model. They retrain weekly and register new model versions in Unity Catalog. They want the endpoint to automatically pick up the newest registered version without manual intervention, while keeping the previous version available for instant rollback. Which approach should they implement?

A.Enable automatic model version detection on the endpoint so it polls the registry every hour and serves the highest version number.
B.Register each weekly model under the same Unity Catalog model name and rely on MLflow's 'latest' stage to route production traffic automatically.
C.Create a scheduled Databricks job that calls the serving endpoint's update-config REST API with the new version ID whenever a training run completes.
D.Configure the endpoint with the Unity Catalog model name and the alias 'champion'; update the alias to the new version after each weekly training job.
AnswerD

Serving endpoints can reference a Unity Catalog model by name plus an alias, and traffic follows whatever version the alias points to. Updating the champion alias after each training job makes the newest version live without editing the endpoint config, and the prior version stays addressable by its version number or a challenger alias for rollback.

Why this answer

Aliases in Unity Catalog give a stable, named pointer to a model version that a serving endpoint can consume. Promoting a freshly validated version means moving the champion alias to that version, which instantly changes production traffic while leaving earlier versions intact for rollback. This removes manual endpoint edits and custom API automation from the weekly promotion path.

Exam trap

The trap here is assuming an endpoint can auto-discover the newest model version, when serving actually binds to an explicit version or alias that must be promoted.

39
MCQhard

You are monitoring a production model deployed on Databricks Model Serving. You notice that the model's predictions have gradually become less accurate over time, likely due to data drift. You need to implement a solution that automatically detects drift and triggers retraining. Which Databricks feature should you use to monitor the model's input data and performance?

A.Databricks SQL dashboards that query the inference table and display accuracy metrics.
B.MLflow Tracking to log prediction results and manually compare distributions over time.
C.Databricks Lakehouse Monitoring for model quality and data drift, configured on the inference table that logs requests and responses.
D.Unity Catalog lineage to track data transformations and identify changes in input data sources.
AnswerC

Databricks Lakehouse Monitoring can be configured on inference tables to track data drift and model performance metrics over time. It provides automated alerts and can trigger retraining workflows. This is the native Databricks solution for monitoring model quality in production, integrating with the lakehouse for scalable analysis.

Why this answer

Databricks Lakehouse Monitoring is the appropriate feature to monitor model input data and performance for drift. It can be configured on inference tables to automatically compute drift metrics and alert when thresholds are exceeded, enabling timely retraining. Other options lack the automated drift detection and integration with production workflows required for this scenario.

Exam trap

The trap here is confusing data lineage or manual tracking with automated drift detection, but only Lakehouse Monitoring provides out-of-the-box statistical monitoring for production models.

40
Multi-Selecthard

You are implementing a CI/CD pipeline for a Databricks ML project using Databricks Repos and Databricks Asset Bundles. You need to ensure that the pipeline promotes code and model artifacts across dev, staging, and prod workspaces consistently. Which TWO practices should you implement to achieve this? (Choose two.)

Select 2 answers
A.Register model versions to a shared Unity Catalog metastore that is accessible from all workspaces, using environment-specific model names or aliases.
B.Store workspace-specific configuration in `databricks.yml` target blocks and use bundle variables for environment differences.
C.Store model artifacts in a separate cloud storage bucket outside Databricks and copy them manually between environments.
D.Use the same Unity Catalog catalog name in all workspaces and rely on workspace-local paths for artifacts.
E.Hardcode the production workspace URL and cluster ID in the notebook code to simplify deployment.
AnswersA, B

A shared Unity Catalog metastore allows model versions to be registered once and referenced across workspaces. By using environment-specific model names or aliases (e.g., `model_dev`, `model_prod`), you maintain separation while enabling promotion. This aligns with Databricks' recommended governance model and ensures artifacts are consistently available to all environments.

Why this answer

Using Databricks Asset Bundles with target blocks and variables, and registering models to a shared Unity Catalog metastore with environment-specific names or aliases, together provide a consistent, governed promotion path. These practices externalize configuration and centralize artifact management, which are core to reliable CI/CD across multiple workspaces.

Exam trap

The trap here is assuming that using identical catalog names or hardcoding workspace details simplifies multi-environment deployment, when it actually breaks isolation and portability.

41
MCQmedium

A machine learning engineer is using MLflow to track experiments and has logged a model with a signature. They now want to register this model in the MLflow Model Registry and promote it to Production. Which MLflow API call should they use to add the model to the registry?

A.mlflow.log_model(model, artifact_path, registered_model_name=name)
B.mlflow.pyfunc.save_model(path, loader_module, data_path, registered_model_name=name)
C.mlflow.register_model(model_uri, name)
D.mlflow.create_registered_model(name)
AnswerC

mlflow.register_model is the correct API to register an existing model artifact from a run into the Model Registry. It takes the model URI (e.g., runs:/<run_id>/model) and the desired model name. This call creates a new model version in the registry, which can then be transitioned to Production.

Why this answer

The mlflow.register_model function is designed to take an existing model URI and add it as a new version to the Model Registry. It is the standard way to register a model after logging. Other options either create empty models or log new models, which do not fit the scenario of promoting an already logged model.

Exam trap

The trap here is confusing the API for creating a registered model container with the API for adding a version from an existing run.

42
MCQhard

A machine learning engineer is setting up a CI/CD pipeline for a model deployed to Databricks Model Serving. The pipeline must automatically update the serving endpoint when a new model version is registered in the MLflow Model Registry and passes a validation job. Which Databricks feature should the engineer use to trigger the update?

A.Databricks Model Serving's built-in auto-update feature, which automatically deploys the latest model version from the registry.
B.MLflow Model Registry webhooks that call a Databricks job to validate and update the endpoint.
C.A Databricks notebook with a while loop that continuously checks the Model Registry for new versions and updates the endpoint.
D.Databricks Jobs with a schedule trigger that polls the Model Registry every hour for new versions.
AnswerB

Registry webhooks can trigger on MODEL_VERSION_CREATED events, invoking a Databricks job that runs validation and then updates the serving endpoint via the REST API. This provides an event-driven, automated pipeline that meets the requirement of updating upon registration and successful validation.

Why this answer

MLflow Model Registry webhooks enable event-driven automation. By configuring a webhook on MODEL_VERSION_CREATED, a Databricks job can be triggered to validate the new version and then update the Model Serving endpoint via the REST API. This integrates validation gating and deployment in a CI/CD pipeline.

Exam trap

The trap here is assuming Model Serving has an auto-update feature or that polling is sufficient, when event-driven webhooks are the intended mechanism for registry-triggered deployments.

43
MCQeasy

A machine learning engineer is setting up a Databricks job to retrain a model every night. The job must use a specific Python library version that is not available in the default Databricks Runtime. The engineer wants to ensure that the retraining job has access to this library without affecting other jobs in the workspace. What is the recommended approach?

A.Modify the global init script to install the library on all clusters in the workspace.
B.Install the library on the job cluster by adding it as a cluster library, and configure the job to use that cluster.
C.Create a custom Docker image with the required library and configure the job cluster to use that image.
D.Install the library using `%pip install` in the first cell of the notebook used by the job.
AnswerB

In Databricks, you can install libraries at the cluster level, making them available to all notebooks and jobs that run on that cluster. For a job, you can define a job cluster with the required library installed. This ensures the library is present for every run without manual installation, and it isolates the library to that cluster, preventing impact on other jobs. This is the recommended and simplest approach for job-specific dependencies.

Why this answer

For a Databricks job that requires a specific library, the best practice is to install the library on the job cluster. Job clusters are dedicated to the job and can be configured with libraries that are installed at cluster startup. This ensures the library is available for every run, isolates dependencies, and avoids impacting other workloads.

It is simpler and more maintainable than custom Docker images or notebook-level installations.

Exam trap

The trap here is thinking that installing a library in a notebook with `%pip install` is sufficient for a scheduled job, but that installation is ephemeral and not isolated, so it fails to provide a persistent, job-specific environment.

44
MCQeasy

Which Databricks feature is specifically designed to prevent data leakage during model training by ensuring feature values are fetched as they existed at a specific point in time?

A.Delta Lake Time Travel.
B.Databricks Feature Store.
C.MLflow Model Registry.
D.Auto Loader.
AnswerB

The Feature Store automatically handles point-in-time joins using specified lookup keys and timestamps. This functionality is essential for preventing data leakage during training, as it ensures that the training dataset only contains information that would have been available at the moment of prediction in a real-world setting.

Why this answer

Point-in-time joins are the core mechanism of the Databricks Feature Store. When training on historical data, it is critical to use features that were available at that time, not their current values. This prevents 'look-ahead bias,' where future information inadvertently leaks into the training set, causing the model to perform artificially well during training but failing in production.

Exam trap

Students frequently confuse standard data caching or cross-validation with point-in-time correctness, missing the specialized purpose of the Databricks Feature Store.

45
MCQhard

You are debugging a model serving issue where the model is failing in production. Which THREE of the following actions should you prioritize to identify the root cause of the failure?

A.Check the Databricks cluster/endpoint logs for errors during the inference request.
B.Review the input data schema and compare it with the training data expectations.
C.Retrain the model from scratch on the full production dataset to clear errors.
D.Examine the model lineage in MLflow to identify the exact code and data used.
E.Delete the model registry and recreate the model versions.
AnswerA, B, D

Logs are the primary source of truth for runtime failures. They will contain stack traces if the model code crashed, memory exhaustion errors, or connectivity issues with external services. Reviewing logs is the fastest way to determine if the issue is environmental or logic-based.

Why this answer

Effective troubleshooting in MLOps relies on systematic isolation of the failure point: infrastructure, data, or model code. Checking system logs, verifying input data schemas, and reviewing model version history are standard diagnostic steps. These actions help determine if the failure is due to a sudden change in incoming data distributions, an infrastructure error, or a regression within the specific model version currently deployed to production.

46
Multi-Selecthard

You are implementing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically run unit tests, train the model, and deploy it to a staging endpoint. Which TWO practices should you follow to ensure the pipeline is reproducible and reliable? (Choose two.)

Select 2 answers
A.Use the latest version of all libraries to ensure compatibility and security patches.
B.Pin all library dependencies to specific versions in a requirements file or conda environment specification.
C.Configure the pipeline to run on a cluster with autoscaling enabled to handle variable workloads.
D.Manually promote the model to staging after reviewing the test results to ensure quality.
E.Store all code, including notebooks and Python modules, in a Git repository and use Databricks Repos to check out the code in the pipeline.
AnswersB, E

Pinning dependencies ensures that the environment is consistent across runs, preventing issues caused by library updates. This is critical for reproducibility in ML pipelines, where even minor version changes can affect model training and inference. Using a requirements file or conda environment specification allows the pipeline to recreate the exact environment.

Why this answer

For a reproducible and reliable CI/CD pipeline, code should be version-controlled and checked out consistently, and dependencies should be pinned to specific versions. These practices ensure that the pipeline runs the same code in the same environment every time. Autoscaling and manual promotion do not address reproducibility, and using latest libraries can introduce instability.

Exam trap

The trap here is assuming that using the latest libraries or manual promotion improves reliability, when in fact they introduce variability and reduce automation.

47
MCQmedium

A fraud detection model is served via a Databricks Model Serving endpoint. The team wants to capture the incoming request payloads and the model's predictions to a Delta table for monitoring and future retraining. Which approach is most appropriate?

A.Enable inference tables on the model serving endpoint to automatically log requests and responses to a Delta table.
B.Modify the model's predict method to write each request and response to a Delta table using the Spark connector.
C.Configure the endpoint to send logs to a cloud storage bucket and schedule a job to load them into Delta.
D.Use MLflow autologging on the serving endpoint to capture inference requests as MLflow runs.
AnswerA

Inference tables are a Databricks Model Serving feature that captures request and response payloads to a Delta table automatically. This provides the exact logging needed for monitoring and retraining without modifying the model code or building custom logging infrastructure.

Why this answer

Databricks Model Serving provides inference tables as a built-in mechanism to log request and response payloads to a Delta table. This satisfies monitoring and retraining data capture without custom code. The other options either add latency, require unsupported dependencies in the serving environment, or confuse training-time autologging with serving-time logging.

Exam trap

The trap here is assuming you must instrument the model code or build a custom pipeline, when Model Serving already offers inference tables for request and response logging.

48
MCQmedium

Your organization needs to automate the deployment of models to real-time serving endpoints. Which service within the Databricks ecosystem handles the hosting and scaling of these endpoints with managed containerization?

A.Databricks Jobs
B.Databricks Model Serving
C.Delta Live Tables
D.Databricks Repos
AnswerB

Model Serving is the dedicated service for hosting models. It manages the server infrastructure, auto-scaling, and deployment of models from the registry. This allows data science teams to deploy models via a single click or API call, ensuring consistent, scalable production performance without manual configuration of servers.

Why this answer

Databricks Model Serving provides a fully managed, scalable infrastructure for deploying models as REST APIs. It handles the underlying container orchestration, allowing developers to focus on the model artifact rather than infrastructure management. This is essential for MLOps, as it simplifies the transition from a registered model in the MLflow registry to a production-ready, highly available endpoint that handles real-time traffic efficiently.

Exam trap

Candidates often confuse Databricks Model Serving with MLflow tracking or model registry features, picking general tracking tools instead of the dedicated endpoint hosting service designed for real-time containerized serving.

49
MCQmedium

Your team is migrating models to Databricks Model Registry. You need to automate the transition of a model version to 'Staging' only after it passes an automated integration test suite in the CI/CD pipeline. Which mechanism should you use to best achieve this?

A.Manually update the model stage using the Databricks UI after the tests complete.
B.Configure a Databricks Job to transition the model based on a scheduled cron trigger.
C.Use the MLflow Python client's transition_model_version_stage method within the CI/CD script.
D.Set the model stage to 'Staging' inside the training notebook using mlflow.register_model.
AnswerC

The 'transition_model_version_stage' method is the standard programmatic way to manage lifecycle transitions in Databricks. By embedding this in the pipeline, the system verifies test results first, then immediately updates the stage, ensuring that only validated models are marked as ready for further staging or production deployment.

Why this answer

Integrating the MLflow client within the CI/CD pipeline allows for programmatic stage transitions using the 'transition_model_version_stage' API. This approach ensures that manual intervention is minimized and models only move forward after passing defined quality gates. Automating this lifecycle phase is a cornerstone of robust MLOps, as it prevents untested models from reaching production environments and ensures consistency across deployment stages.

50
MCQmedium

When deploying a model to a production environment, why is it recommended to use a 'Model Signature'?

A.To increase the inference speed of the model.
B.To define the input and output schema for automatic validation.
C.To automatically retrain the model if the input data changes.
D.To encrypt the model artifacts in the registry.
AnswerB

The signature provides a contract for the model, detailing the expected data types and shapes. The serving infrastructure uses this contract to validate every API request, catching mismatches before they reach the model code. This prevents crashes and provides clear error messages when invalid data is submitted.

Why this answer

A model signature defines the schema of the inputs and outputs. By enforcing this, the serving endpoint can automatically validate incoming requests and reject those that do not match the expected format. This prevents runtime errors and unexpected behavior, serving as a critical safety feature in automated production pipelines where data quality might fluctuate, ensuring the model always receives the correct data structure.

Exam trap

Candidates often confuse model signatures with performance metrics or hyperparameters, failing to realize that signatures specifically define the expected input and output data schemas.

51
MCQeasy

A machine learning engineer needs to automate retraining of a model whenever new data lands in a Delta table. The retraining must run on a schedule, use a specific cluster configuration, and send an email alert on failure. Which Databricks feature should be used to orchestrate this workflow?

A.Databricks Jobs with a notebook task, a schedule trigger, and an email notification on failure.
B.An MLflow Project run triggered by a file arrival event in DBFS using a filesystem watcher.
C.A Delta Live Tables pipeline with a continuously running mode that trains the model in a streaming table.
D.A Databricks SQL dashboard with a scheduled refresh that calls a stored procedure to retrain the model.
AnswerA

Databricks Jobs provide scheduling, cluster specification, and notification configuration in one place. A notebook task can read the Delta table, retrain, and register the model. Schedule triggers and failure notifications are built-in, directly meeting all stated requirements without external orchestration tools.

Why this answer

Databricks Jobs is the native orchestration service for scheduled notebooks and scripts. It supports specifying cluster configuration, scheduling, and failure notifications. The other options either lack scheduling, cannot run training code, or are designed for data pipelines rather than ML workflows, so they do not satisfy the combined requirements of schedule, cluster control, and alerting.

Exam trap

The trap here is reaching for specialized ML features like MLflow Projects or Delta Live Tables when the requirements are basic scheduling, cluster control, and alerting that Databricks Jobs already provides.

52
Multi-Selectmedium

You are responsible for monitoring a critical model deployed to Databricks Model Serving. You need to detect data drift and model performance degradation. Which TWO of the following actions should you take? (Choose two.)

Select 2 answers
A.Deploy the model to a different workspace to isolate production traffic.
B.Set up a Databricks SQL dashboard that queries the inference table and compares feature distributions to a baseline.
C.Enable inference tables on the model endpoint to log request and response data.
D.Use MLflow to log the model's training metrics and compare them to production metrics manually.
E.Configure the endpoint to use a smaller instance type to reduce cost.
AnswersB, C

Databricks SQL dashboards can query inference tables to compute statistics and compare them against baseline distributions. This allows you to visualize drift and set up alerts. By leveraging SQL, you can create scheduled queries that detect significant deviations and trigger notifications, providing an effective monitoring solution without additional infrastructure.

Why this answer

Inference tables capture the necessary request and response data for monitoring, while Databricks SQL dashboards enable analysis and alerting on that data. Together, they provide a robust solution for detecting data drift and model performance degradation. Other actions, such as changing instance types or isolating workspaces, do not address the monitoring requirements.

Exam trap

The trap here is focusing on infrastructure changes or manual comparisons instead of leveraging automated data capture and analysis tools designed for monitoring.

53
MCQmedium

You are responsible for deploying a machine learning model to a Databricks Model Serving endpoint. The model must be updated frequently with new versions. You want to ensure that the endpoint remains available during updates and that traffic can be shifted gradually to the new version. Which deployment strategy should you use?

A.Rolling update by stopping the endpoint, updating the model, and restarting the endpoint.
B.Canary deployment by configuring the serving endpoint with multiple model versions and specifying a traffic split percentage.
C.Shadow deployment by duplicating requests to a new endpoint without affecting production traffic.
D.Blue-green deployment by creating a new endpoint and switching traffic via DNS.
AnswerB

Databricks Model Serving supports serving multiple model versions on a single endpoint with configurable traffic splits. This enables canary deployments where a small percentage of traffic goes to the new version, allowing gradual rollout and rollback if issues arise. This is the native and recommended strategy for safe updates.

Why this answer

Databricks Model Serving allows you to configure multiple model versions on a single endpoint and set traffic split percentages. This enables canary deployments, where you can gradually shift traffic to a new version while monitoring performance, ensuring high availability and easy rollback. Other strategies either cause downtime or are not natively supported.

Exam trap

The trap here is assuming that classic deployment strategies like blue-green or rolling updates apply directly, but Databricks Model Serving provides built-in traffic splitting for canary deployments.

54
MCQhard

An organization needs to implement a robust CI/CD strategy for their Databricks ML models. Which TWO of the following practices are recommended to ensure reliable model deployment?

A.Perform manual model registration directly from the production workspace notebooks.
B.Use Databricks Repos to synchronize code across development, staging, and production environments.
C.Execute all training runs directly in the production workspace for maximum accuracy.
D.Implement automated model validation tests before promoting a model to the 'Production' stage.
E.Store all raw data directly inside the MLflow model artifact for portability.
AnswerB, D

Databricks Repos allows teams to manage code using Git, which is essential for CI/CD. Synchronizing code across environments ensures that the exact same logic tested in staging is deployed in production, minimizing the risk of 'it works on my machine' errors and facilitating smooth, reproducible deployments.

Why this answer

Implementing automated testing and environment isolation are cornerstones of MLOps. Automated unit and integration tests catch regressions early, while separating development, staging, and production workspaces prevents accidental interference with live systems. Using Git-based workflows via Databricks Repos ensures code traceability, while the Model Registry acts as the gatekeeper for model promotion, ensuring only validated models reach production environments.

Exam trap

Candidates often pick options involving manual model retries or simple notebook exports, failing to recognize that CI/CD requires Git integration and automated gatekeeping via the Model Registry.

55
MCQeasy

In the context of Databricks MLOps, what is the primary purpose of a 'Staging' environment in the Model Registry?

A.To act as a long-term backup for model version history.
B.To allow for final validation and testing before production deployment.
C.To automatically convert models into a different serving format.
D.To increase the model's training accuracy.
AnswerB

Staging acts as a testing sandbox where teams can run integration tests or shadow deployments against live data. This ensures that the model is fully vetted for stability and performance in a controlled environment similar to production, mitigating the risk of issues when the model goes live.

Why this answer

The Staging environment serves as a pre-production gate where models undergo final validation, such as integration testing and UAT (User Acceptance Testing), before being promoted to Production. This ensures that only models that meet performance, safety, and operational standards are exposed to live production traffic, reducing the risk of catastrophic failures in the real-world deployment environment.

Exam trap

Test-takers frequently confuse the 'Staging' environment with a live testing zone for end-users, missing its true purpose as an automated pre-production validation gate.

56
MCQmedium

A machine learning engineer needs to track model experiments and ensure that all training parameters and metrics are captured in a reproducible way. What is the best practice for using MLflow within Databricks notebooks?

A.Manually log metrics to a shared Excel file stored on DBFS after the run completes.
B.Use MLflow with 'mlflow.start_run()' and 'mlflow.autolog()' to ensure all runs are tracked automatically.
C.Only log the final evaluation metric at the end of the training loop to minimize storage.
D.Print the parameters to the notebook output cell and rely on the notebook's revision history.
AnswerB

Using 'mlflow.start_run()' within a context manager ensures that all tracking information is cleaned up correctly, even if the code fails. 'mlflow.autolog()' reduces boilerplate code by capturing common model parameters and metrics automatically, significantly improving the consistency and efficiency of the experiment tracking process across the organization.

Why this answer

Using 'mlflow.autolog()' or explicit 'mlflow.log_params()' within a context manager ensures that every run is isolated and documented. This practice allows for easy comparison of different hyperparameter configurations, preventing data loss and providing a clear lineage from the raw data to the final model artifact, which is crucial for compliance and debugging in enterprise ML environments.

Exam trap

Candidates often mention manual logging of every parameter individually, missing that 'mlflow.autolog()' is the best practice for capturing comprehensive run data with minimal code overhead.

57
MCQhard

A team runs a nightly Databricks job that retrains a demand-forecasting model and registers a new version in the MLflow Model Registry. Compliance requires that the model version used for scoring in production be immutably identified and that any subsequent retraining not silently change what production serves. Which practice best satisfies this requirement?

A.Register each retrained model under a new registered model name that encodes the training date, and update the scoring job to read the newest name each night.
B.Store the model artifact path from the training run in a Delta table and have the scoring job read the path directly, bypassing the Model Registry entirely.
C.Have the scoring job reference the model by its registered model name only, so that the latest version is always used after each nightly run.
D.Assign a stage or alias to the specific model version and have the scoring job reference the model by that stage or alias, updating it only through a controlled promotion step.
AnswerD

Stages and aliases provide an indirection pointer to a specific model version. By promoting only through a controlled step, the scoring job always resolves to the exact version that was validated, and retraining that registers new versions does not alter production until an explicit promotion changes the pointer, satisfying immutability and traceability.

Why this answer

Pinning production to a specific model version through a stage or alias gives an auditable pointer that only changes via an explicit promotion. Referencing the model name alone, inventing new names per run, or reading raw artifact paths all allow production to drift or lose lineage, so they fail the immutability and traceability requirement.

Exam trap

The trap here is believing that referencing the registered model name always serves a fixed artifact, when it actually resolves to whatever version the current stage or alias points to.

58
MCQmedium

Your team is using MLflow Model Registry to manage a model that predicts customer churn. A new version has been registered and passed validation. You need to transition this version to the 'Production' stage and ensure that all downstream scoring jobs automatically use it. What should you do?

A.Keep the model version in 'Staging' and update the downstream jobs to reference the model by stage 'Staging'.
B.Transition the model version to 'Production' and update the downstream jobs to reference the model by stage 'Production'.
C.Archive the current production model and register the new version with a new model name.
D.Transition the model version to 'Production' and update the downstream jobs to reference the model by version number.
AnswerB

Transitioning to 'Production' and having jobs reference the stage ensures that when a new version is promoted, the jobs automatically pick it up. This is the intended workflow for stage-based model deployment in MLflow Model Registry, enabling seamless updates without code changes.

Why this answer

Transitioning the model version to the 'Production' stage and having downstream jobs reference the model by stage 'Production' is the correct approach. This allows you to promote new versions to production without modifying job code, ensuring that all consumers automatically use the latest approved model. It also maintains a clear audit trail and governance.

Exam trap

The trap here is confusing version-based referencing with stage-based referencing; only stage-based referencing enables automatic updates when a new version is promoted.

59
MCQeasy

A team wants to deploy a scikit-learn model to a real-time REST endpoint on Databricks. They have logged the model with MLflow and registered it in Unity Catalog. Which method should they use to create the serving endpoint?

A.Use the Databricks Model Serving UI or REST API to create an endpoint and select the registered model version.
B.Package the model into a Docker image and deploy it to a Kubernetes cluster managed by Databricks.
C.Write a Databricks job that loads the model and exposes it via a Flask app on a driver node.
D.Use MLflow's `mlflow models serve` command on a Databricks cluster to start a local REST server.
AnswerA

Databricks Model Serving integrates directly with Unity Catalog and MLflow. You create an endpoint via the Serving UI or the REST API, specifying the full model name and version. The service automatically provisions the necessary infrastructure and builds a container with the model's dependencies, making this the standard and supported deployment path.

Why this answer

Databricks Model Serving is the managed solution for deploying registered models as REST endpoints. It handles infrastructure, scaling, and dependency management automatically. Using the UI or REST API to create an endpoint from a Unity Catalog model version is the correct and supported approach, unlike manual containerization or local serving.

Exam trap

The trap here is confusing local MLflow serving commands or custom Flask apps with the managed Databricks Model Serving feature.

60
MCQhard

You are implementing a CI/CD pipeline for a model. You want to automate unit testing of the model's inference performance. Which approach is best suited for Databricks?

A.Manually inspect performance metrics in MLflow UI before every deployment.
B.Trigger a notebook to run inference on a validation set and assert metrics.
C.Deploy the model to staging and wait for user feedback.
D.Rely on the Model Registry to automatically validate model performance.
AnswerB

Automating validation via notebooks allows for programmatic assertion of metrics like accuracy or latency. This creates a reliable quality gate that runs as part of the deployment pipeline. If the model fails the assertions, the pipeline halts, preventing the promotion of a sub-par model to production.

Why this answer

Using a dedicated job workflow that triggers a notebook to run inference on a hold-out test dataset is the standard Databricks approach. By comparing metrics against defined thresholds before deployment, you ensure quality. This automated gate prevents poor-performing models from reaching production, which is a critical step in a mature MLOps lifecycle to maintain system stability and model accuracy.

Exam trap

Candidates often suggest using external CI/CD tools like Jenkins or GitHub Actions to perform model testing, forgetting that Databricks jobs can natively run notebooks to validate inference performance.

61
MCQmedium

Which THREE of the following are key responsibilities of an MLOps engineer when managing a model lifecycle on Databricks?

A.Automating the model registration and promotion process in CI/CD.
B.Writing the core machine learning algorithms for every project.
C.Monitoring the production health and data drift of deployed models.
D.Managing access control and governance of model artifacts.
E.Manually testing every model in the production environment.
AnswerA, C, D

Automation is at the heart of MLOps. By scripting the model registry interactions within CI/CD pipelines, engineers ensure that models move through stages consistently, preventing manual errors and ensuring that the promotion process is documented and repeatable across the organization's different deployment environments.

Why this answer

MLOps engineers act as the bridge between data science and production operations. Their core duties involve automating the lifecycle, ensuring that data quality is monitored, and maintaining the infrastructure security and governance required to serve models reliably. These activities ensure that machine learning remains a stable and repeatable business function rather than an ad-hoc, error-prone research endeavor.

Exam trap

Candidates frequently select ad-hoc data science tasks like hyperparameter tuning or feature engineering as primary MLOps responsibilities, ignoring governance, CI/CD promotion automation, and production monitoring.

62
Multi-Selectmedium

A financial institution is using Databricks to build and deploy a credit risk model. The model must comply with regulations that require full auditability of the model's lineage, including data sources, transformations, and training parameters. The team uses MLflow for tracking and the Feature Store for feature management. Which TWO of the following practices are essential to meet the auditability requirements? (Choose two.)

Select 2 answers
A.Use Databricks Feature Store to create feature tables and log the training set with the model, including feature lookup keys.
B.Log all input data versions and feature table versions used in training as MLflow tags or artifacts.
C.Enable MLflow autologging for all supported frameworks to capture parameters and metrics automatically.
D.Manually document the data sources and transformations in a separate spreadsheet for each model version.
E.Store all training code in a Git repository and commit after each run.
AnswersA, B

The Feature Store allows you to create feature tables with versioning and lineage. When you train a model using features from the Feature Store and log the model with mlflow.pyfunc.log_model, it automatically records the feature lookup keys and the feature table versions. This creates a direct link between the model and the features used, satisfying audit requirements for feature lineage.

Why this answer

To meet auditability requirements, it is essential to automatically capture data and feature lineage. Logging data versions and feature table versions in MLflow provides a record of the exact inputs used. Using the Feature Store to log the training set with the model ensures that feature lineage is preserved.

Together, these practices create a transparent and verifiable chain from data to model, which is necessary for regulatory compliance.

Exam trap

The trap here is thinking that code versioning or autologging alone suffices for auditability, when the critical missing piece is data and feature lineage.

63
Multi-Selecthard

You are building a CI/CD pipeline that must promote an MLflow model version from Staging to Production in Databricks only after automated validation. The pipeline runs in a service principal context. Which two actions are required to implement this safely and repeatably? (Choose two.)

Select 2 answers
A.Grant the service principal CAN_MANAGE_PRODUCTION (or CAN_MANAGE) on the registered model so it can transition versions into Production.
B.Configure the pipeline to call the MLflow Model Registry transition-stage API after validation succeeds, using the model name and version.
C.Enable automatic model version archiving so that older Production versions are removed before the new version is promoted.
D.Require the service principal to use a personal access token tied to a human user so that audit logs show a named approver.
E.Store the model artifacts in a Unity Catalog volume and reference the volume path when registering the model version.
AnswersA, B

Stage transitions in MLflow Model Registry are governed by model-level permissions. A service principal running the pipeline needs at least CAN_MANAGE_PRODUCTION or CAN_MANAGE to move a version into Production. Without this grant, the API call to transition the stage will fail with a permission error, making the pipeline non-repeatable. This is a required, concrete step for automated promotion.

Why this answer

Safe automated promotion needs two things: permission for the automation identity to change stages, and a programmatic call to perform the transition after validation. Granting the service principal CAN_MANAGE_PRODUCTION or CAN_MANAGE satisfies the first, and invoking the transition-stage API after validation satisfies the second. Archiving policies, artifact storage changes, and human tokens are neither required nor advisable for this pipeline.

Exam trap

The trap here is focusing on artifact storage or archiving policies, when the required elements are the service principal's model permission and the programmatic stage transition call.

64
Multi-Selecthard

You are implementing a CI/CD pipeline for ML models on Databricks. The pipeline must automatically retrain, validate, and promote models to Production in MLflow Model Registry. Which TWO practices are essential for maintaining reproducibility and governance in this automated workflow? (Choose two.)

Select 2 answers
A.Configure the pipeline to automatically transition any model version with accuracy above a fixed threshold directly to Production without human review.
B.Disable model versioning in the registry and instead use run IDs as the sole identifier for production models.
C.Use MLflow Model Registry stage transitions with recorded descriptions and tags that identify the pipeline run and approver for each promotion.
D.Log the Git commit SHA and environment specification with each MLflow run so the exact code and dependencies can be reproduced.
E.Store model artifacts in a separate cloud storage bucket outside the MLflow tracking server to avoid coupling with the registry.
AnswersC, D

Recording descriptions and tags on stage transitions provides an audit trail linking each promotion to a specific pipeline run and approver. This supports governance by making it possible to trace who or what promoted a model and why, which is essential for compliance and incident investigation.

Why this answer

Reproducibility and governance in automated pipelines require capturing the exact code and environment for each run, and maintaining an auditable record of stage transitions with descriptions and tags. Automatic promotion without review and decoupling artifacts or disabling versioning all weaken traceability and control, making them unsuitable for a governed CI/CD workflow.

Exam trap

The trap here is equating automated promotion with governance, when governance actually requires traceability of code, environment, and promotion decisions.

65
MCQmedium

A team wants to track model performance over time to detect drift without writing custom monitoring infrastructure. What is the most efficient Databricks tool for this?

A.MLflow Tracking APIs.
B.Databricks Model Monitoring (DMM).
C.Delta Lake Change Data Feed.
D.Databricks SQL Dashboards.
AnswerB

DMM is the native solution for monitoring production models. It automates drift detection, schema validation, and performance tracking, providing ready-made dashboards and alerts. It eliminates the need to build custom monitoring infrastructure, making it the most efficient choice for teams wanting to maintain high-quality models in production.

Why this answer

Databricks Model Monitoring (DMM) provides built-in capabilities to track model performance, data drift, and input quality. By automating these tasks, teams avoid the technical debt of building and maintaining custom monitoring solutions. DMM integrates seamlessly with the Model Registry, allowing for automated alerts and reports that keep stakeholders informed about model health in production environments.

Exam trap

Test-takers might suggest writing custom Spark jobs or MLflow logging scripts, unaware that Databricks Model Monitoring provides native, out-of-the-box tracking.

66
MCQhard

Refer to the exhibit. A data scientist attempted to register a new model version but received the error shown in the exhibit. Which step should be taken to resolve this issue?

A.Upgrade the user's Databricks entitlement to 'Workspace Admin'.
B.Update the MLflow experiment tracking URI to point to an external database.
C.Ask an administrator to grant 'CAN_MANAGE' or 'CAN_EDIT' permissions on the model in the MLflow UI.
D.Re-run the training notebook using the cluster owner's credentials.
AnswerC

Databricks implements model-specific permissions. To register or modify model versions, the user must have the appropriate permission level assigned to that specific model object. This follows standard security best practices by limiting access to authorized users without granting excessive, unnecessary privileges across the entire workspace or environment.

Why this answer

The error indicates a lack of necessary permissions to perform operations on the MLflow Model Registry in the Databricks workspace. Access control in Databricks is granular; even if a user can write code, they need specific registry-level permissions to promote or register models. This is a common MLOps governance issue where security policies must be correctly balanced with developer autonomy.

Exam trap

Candidates often suggest re-training the model or checking the code, overlooking that this is an IAM/RBAC issue within the Databricks Workspace specific to the Model Registry.

67
Multi-Selecthard

You are implementing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically retrain the model when new data arrives, validate its performance, and promote it to production if it meets quality thresholds. Which TWO of the following steps are essential to include in the pipeline to ensure safe and automated deployment? (Choose two.)

Select 2 answers
A.Implement automated tests that evaluate the model on a holdout dataset and compare metrics against a baseline before promotion.
B.Manually review the model's performance metrics in a notebook before allowing the pipeline to proceed.
C.Store the model artifacts in a Git repository alongside the code to ensure versioning.
D.Configure the pipeline to directly overwrite the production model endpoint with the newly trained model without validation to reduce latency.
E.Use MLflow Model Registry to register the model and transition it to 'Staging' for validation, then to 'Production' after passing tests.
AnswersA, E

Automated tests on a holdout dataset verify that the model meets performance thresholds before promotion. Comparing against a baseline ensures that the new model does not regress. This step is essential for safe automation, as it gates deployment on objective criteria and prevents degradation in production.

Why this answer

The essential steps for an automated CI/CD pipeline include registering the model in MLflow Model Registry with stage transitions and implementing automated tests on a holdout dataset to validate performance against a baseline. These steps ensure that only models meeting quality thresholds are promoted, enabling safe automation. Other steps like manual review or storing artifacts in Git are not suitable for automated pipelines.

Exam trap

The trap here is thinking that manual review or direct overwrite can be part of an automated pipeline, but automation requires objective, reproducible validation steps without human intervention.

68
Multi-Selecthard

A fraud-detection model is served via a Databricks Model Serving endpoint. Compliance requires that every prediction request be traceable to the exact model artifact that produced it and that the model's inputs be auditable for drift analysis. Which two actions should you take to satisfy these requirements? (Choose two.)

Select 2 answers
A.Set the endpoint's scale-to-zero behavior to disabled so logs are never lost during cold starts.
B.Configure the endpoint to use a custom container image that writes each request to an external S3 bucket.
C.Enable inference tables on the endpoint so requests and responses are automatically logged to a Delta table.
D.Tag the registered model version in Unity Catalog with the deployment timestamp and serving endpoint name.
E.Enable automatic model version capture on the endpoint so the served model version is recorded with request logs.
AnswersC, E

Inference tables capture the payloads sent to and returned from a Model Serving endpoint and persist them to a Unity Catalog Delta table. This gives an auditable record of the model inputs for drift analysis and links each request to the served model version, satisfying both traceability and auditability requirements without custom logging code.

Why this answer

Inference tables persist endpoint request and response payloads to a governed Delta table, and enabling model version capture records which served model version produced each prediction. Together they provide the per-request audit trail and the artifact-level traceability the compliance scenario demands, without custom containers or manual tagging that cannot capture payloads.

Exam trap

The trap here is assuming that tagging the model version or keeping replicas warm is enough for auditability, when only payload logging and version capture actually record each prediction.

69
MCQhard

An ML engineering team maintains a feature table in Databricks Feature Store that is populated by a nightly batch job. The same features feed both an offline training pipeline and a Model Serving endpoint. After a schema change adds two columns to the source Delta table, the endpoint begins returning errors during online lookup. The team confirms the offline training pipeline still works. Which change most likely restores the endpoint?

A.Publish the updated feature table to the online store so the online table schema matches the offline table, then redeploy the model that consumes those features.
B.Add the new columns to the model signature and register a new model version, leaving the online store untouched.
C.Recreate the serving endpoint with a larger workload size so the added columns can be cached in memory during lookups.
D.Switch the endpoint to read features directly from the offline Delta table instead of the online store.
AnswerA

The online store is a separate materialized copy optimized for low-latency lookup, and it does not automatically inherit schema changes made to the offline Delta table. Publishing again refreshes the online table to include the new columns so the serving lookup matches what the model expects. Redeploying ensures the model signature aligns with the updated feature schema.

Why this answer

Feature Store maintains two representations: the offline Delta table for training and an online table for low-latency serving. Schema changes flow to the offline table automatically but must be republished to refresh the online store. Because the endpoint queries the online store, its stale schema causes lookup errors even though training still succeeds.

Republishing and redeploying restores alignment.

Exam trap

The trap here is assuming the online store automatically tracks schema changes applied to the offline feature table.

70
MCQeasy

A data scientist registers a new model version to Unity Catalog and wants to promote it to Production after validation. Which action accomplishes this in the current Databricks recommendation?

A.Set the model version's stage to Production using the MLflow Model Registry stage transition API.
B.Assign the alias 'champion' to the validated model version using the MLflow client or Catalog Explorer.
C.Add a tag named 'stage' with the value 'Production' to the model version and restart the serving endpoint.
D.Create a new registered model with the same name plus a '-prod' suffix and copy the artifacts into it.
AnswerB

Databricks recommends using aliases such as 'champion' or 'Production' to mark the version that serving endpoints and jobs should consume. Assigning the alias to the validated version promotes it without mutating the version itself, and endpoints configured to serve that alias pick up the change.

Why this answer

In Unity Catalog, model promotion is done with aliases rather than the deprecated stages. Assigning an alias such as 'champion' to a validated version marks it as the one that serving endpoints and jobs should consume, and endpoints referencing that alias automatically serve the newly promoted version without changing their configuration.

Exam trap

The trap here is assuming that MLflow stages still govern promotion in Unity Catalog, when aliases are the supported mechanism.

71
MCQhard

A financial institution has deployed a credit risk model to a Databricks Model Serving endpoint. The model was trained on data that includes sensitive customer attributes. The compliance team requires that all predictions be explainable and that the model's decisions can be audited. The data science team wants to use SHAP (SHapley Additive exPlanations) to generate explanations for each prediction. Which approach should they take to integrate SHAP with the serving endpoint while maintaining low latency?

A.Enable automatic logging of SHAP values by setting the `log_explainer` parameter in the MLflow model signature.
B.Use the `shap.Explainer` within a custom PyFunc model that computes SHAP values on the fly and returns them alongside predictions.
C.Deploy a separate endpoint that runs SHAP explanations asynchronously and return a job ID that the client can poll for results.
D.Precompute SHAP values for all possible input combinations and store them in a Delta table for lookup during serving.
AnswerB

Wrapping the model in a custom PyFunc that computes SHAP values at inference time allows explanations to be generated for each request. While this adds latency, it can be optimized by using efficient SHAP implementations and limiting the number of background samples. This approach provides real-time explanations and can be deployed to Model Serving. It meets the requirement for per-prediction explainability and auditability, though latency must be managed.

Why this answer

Integrating SHAP into a custom PyFunc model allows the serving endpoint to return both predictions and explanations in real time. This satisfies the compliance need for per-prediction explainability and auditability. While it introduces some latency, it can be optimized with efficient SHAP algorithms and careful resource allocation, making it suitable for real-time serving in regulated environments.

Exam trap

The trap here is assuming that SHAP explanations can be precomputed or automatically logged, when in reality they must be computed at inference time for each request to provide accurate, real-time explanations.

72
MCQeasy

Which of the following describes the 'Gold' layer in the Medallion Architecture, and why is it important for machine learning?

A.It contains raw data ingested directly from external sources.
B.It contains aggregated, business-level data ready for model training.
C.It is used to store model artifacts and logs.
D.It is where data scientists perform exploratory data analysis.
AnswerB

Gold tables provide clean, validated data that matches business requirements. This makes them the ideal source for training models. By building models on Gold data, scientists ensure that their features are based on reliable information, significantly reducing the probability of errors caused by poor data quality in production pipelines.

Why this answer

The Gold layer contains refined, business-level data that is ready for consumption. In MLOps, it is the standard source for training data. By using curated, high-quality data from the Gold layer, data scientists avoid the noise and inconsistencies of raw data, leading to more robust models and faster iteration times, as they spend less time on manual data cleaning and validation.

Exam trap

Candidates often confuse 'Gold' data with 'Silver' data, failing to realize that Gold specifically implies business-level aggregation, which is the necessary state for final model training inputs.

73
MCQmedium

Your organization requires that all models deployed to production undergo a drift detection check. Which approach is most effective for monitoring model performance in Databricks?

A.Write a custom Spark job to compare the training table with inference logs every hour.
B.Enable Databricks Lakehouse Monitoring on the inference Delta tables to track drift automatically.
C.Configure the Model Registry to automatically retrain the model whenever performance dips below a threshold.
D.Rely on end-user feedback to manually flag when the model performance seems degraded.
AnswerB

Lakehouse Monitoring offers native, integrated drift detection that works directly with Unity Catalog and Delta tables. It provides out-of-the-box dashboards and automated alerts, reducing the need for custom code and ensuring that drift detection is consistently applied across all production models in the environment.

Why this answer

Monitoring drift requires comparing inference data distributions against the training baseline. Databricks Lakehouse Monitoring provides a managed service that automatically detects feature and prediction drift. By leveraging Delta tables as the source of truth, the monitoring service can compute statistics periodically and trigger alerts, ensuring that any degradation in model performance is identified quickly, allowing for proactive retraining or model rollback strategies.

Exam trap

Candidates often suggest building custom monitoring dashboards using SQL or Python code, ignoring that Databricks Lakehouse Monitoring is the native, automated tool specifically built for drift detection.

74
MCQmedium

A team is building a feature store in Databricks. They need to ensure that training data and inference data are consistent to avoid training-serving skew. What is the primary benefit of using the Databricks Feature Store in this context?

A.It automatically scales the underlying cluster during training runs.
B.It provides a unified interface for computing features, ensuring consistency between training and inference.
C.It forces data scientists to use only Python for feature engineering.
D.It eliminates the need for any data cleaning or preprocessing steps.
AnswerB

Consistency is achieved by using the same feature definitions and logic for both offline training datasets and online inference lookups. The Feature Store acts as the bridge that guarantees the data fed into the model during training is identical to what the model receives during real-time serving.

Why this answer

The Databricks Feature Store enables the reuse of feature pipelines, ensuring that the exact same transformation logic is applied during both training and inference. By decoupling feature generation from model training, it creates a 'single source of truth' for features, effectively eliminating training-serving skew and improving the reliability of models in production. This is a vital component for maintaining feature consistency across distributed teams.

Exam trap

Candidates often assume the Feature Store is primarily for performance optimization or storage reduction rather than its core purpose of ensuring consistent transformation logic between training and serving.

75
MCQhard

A company uses Databricks Model Serving to host a real-time model. They need to perform A/B testing between two model versions, sending 10% of traffic to a new version and 90% to the current version. Which approach should they use?

A.Log both models as a single MLflow model that internally routes based on a random number.
B.Use MLflow Model Registry stages to mark one version as Production and the other as Staging, then rely on the serving endpoint to automatically split traffic.
C.Deploy both model versions to the same endpoint and use the traffic split feature in Model Serving.
D.Create two separate endpoints, one for each version, and use an external load balancer to split traffic.
AnswerC

Databricks Model Serving supports serving multiple model versions on a single endpoint with configurable traffic percentages. This allows seamless A/B testing by routing a percentage of requests to each version. The feature is built into the endpoint configuration, enabling easy adjustments and monitoring.

Why this answer

Databricks Model Serving allows multiple model versions on one endpoint with specified traffic percentages, making A/B testing straightforward. This native feature handles routing without external components. Other approaches require additional infrastructure or custom logic, which are unnecessary and less efficient.

Exam trap

The trap here is assuming that MLflow stages automatically control traffic splitting, when in fact traffic split must be explicitly configured on the serving endpoint.

Page 1 of 2 · 130 questions totalNext →

Ready to test yourself?

Try a timed practice session using only ML Ops questions.