Courseiva

CCNA ML Ops Questions

55 of 130 questions · Page 2/2 · ML Ops · Answers revealed

76
MCQmedium

Your team uses Databricks Model Serving to host a production model. You need to ensure zero-downtime updates while maintaining the ability to revert to the previous version instantly if performance degrades. Which deployment strategy should you implement?

A.Directly update the existing model serving endpoint to the new version.
B.Use a Blue-Green deployment strategy with traffic splitting.
C.Delete the current serving endpoint and create a new one with the updated model.
D.Increase the instance count of the existing endpoint before updating.
AnswerB

Blue-Green deployment enables side-by-side versions, allowing for controlled traffic shifting. This methodology provides a seamless transition path and an instantaneous rollback capability, which are essential requirements for maintaining service reliability and managing risk in high-stakes production machine learning environments.

Why this answer

Blue-green deployment allows you to maintain a production endpoint while simultaneously deploying a new version. By using weighted traffic splitting, you can shift traffic gradually and monitor performance metrics. If the new version exhibits anomalies, you can instantly revert traffic to the stable version, ensuring high availability and minimizing the impact of potential production failures, which is critical for robust MLOps workflows.

Exam trap

Candidates often confuse 'Blue-Green' with 'A/B testing'. While they share traffic splitting, Blue-Green specifically focuses on seamless cutovers and instant rollbacks for deployment reliability, not just statistical comparison.

77
MCQmedium

A team is transitioning from manual experimentation to a formal MLOps pipeline. Which component should be prioritized to ensure that data used during training is reproducible?

A.Store all training data as CSV files in DBFS with unique filenames for each run.
B.Use Delta Lake time travel to query the training data at the specific timestamp of the model training run.
C.Ask the data engineering team to copy the data to a private folder for every run.
D.Only rely on the model performance metrics to infer the quality of the training data.
AnswerB

Delta Lake time travel allows for precise data versioning. By referencing a specific timestamp or version, the pipeline ensures that the exact training set used in a previous experiment can be perfectly reproduced, which is a fundamental requirement for reliable and compliant MLOps workflows.

Why this answer

Delta Lake's time-travel capability is the standard for ensuring data reproducibility in MLOps. By using 'VERSION AS OF' or 'TIMESTAMP AS OF', you can query the exact state of the training dataset at any point in time. This is critical for auditing, model retraining, and debugging, as it eliminates the ambiguity of changes made to the underlying data sources over time.

Exam trap

Candidates often assume saving a standard CSV file or taking a static snapshot is sufficient, forgetting about Delta Lake's native time-travel features.

78
MCQmedium

Your team is migrating models to Databricks Model Registry. You need to ensure that production models are only deployed after passing specific validation tests. Which feature best facilitates this automated governance?

A.Using MLflow Experiments to log model artifacts.
B.Hard-coding model versions in inference notebooks.
C.Configuring Model Registry webhooks for status transitions.
D.Executing manual SQL queries against the MLflow backend.
AnswerC

Webhooks allow you to register external endpoints that are triggered when specific events occur, such as a model version transition to 'Production'. This enables automated CI/CD integration, ensuring that only verified models move forward, which is a foundational requirement for robust, automated MLOps governance workflows.

Why this answer

Model Registry webhooks are the critical component for automating governance. By triggering external CI/CD pipelines (like Jenkins or GitHub Actions) upon model version transitions (e.g., 'Pending' to 'Production'), organizations enforce quality gates. This ensures that only models validated by automated unit tests or performance benchmarks are promoted, reducing manual intervention and minimizing the risk of deploying underperforming models into production environments.

Exam trap

Candidates often assume manual review processes or standard notification emails can enforce automated governance, ignoring the role of Model Registry webhooks in triggering CI/CD pipelines.

79
Multi-Selectmedium

Which TWO of the following are primary benefits of using the MLflow Model Registry in Databricks?

Select 2 answers
A.Automated training of models.
B.Lifecycle state tracking (e.g., Staging, Production).
C.Centralized versioning and lineage.
D.Real-time monitoring of inference latency.
E.Automatic data cleaning for training sets.
AnswersB, C

Tracking the state of a model version is a core feature of the registry. It allows teams to clearly define which models are ready for production, which are being tested, and which have been deprecated, ensuring that deployment pipelines only pull models from the appropriate, authorized lifecycle stages.

Why this answer

The Model Registry centralizes model management, providing a single source of truth for model lifecycle stages and version history. This is vital for MLOps, as it ensures that teams can track which models are in production, who approved them, and how they perform over time, enabling consistent deployment workflows and improved auditability for regulatory compliance and internal governance.

Exam trap

Candidates sometimes select 'model training' or 'artifact storage' as primary benefits, confusing general MLflow features with the specific governance and lifecycle management provided by the Model Registry.

80
MCQeasy

A team notices that a deployed Model Serving endpoint occasionally returns stale predictions for a subset of customers shortly after the nightly feature refresh completes. The offline feature table is confirmed current. Which is the most likely explanation?

A.The offline feature table's Delta transaction log has not been checkpointed, so readers see an older snapshot.
B.The online feature store was not republished after the nightly refresh, so lookups return the previous feature values for those customers.
C.The model version serving traffic is older than the latest registered version, so predictions lag behind the current model.
D.The endpoint's autoscaling added replicas that cached outdated feature values during scale-out.
AnswerB

The online store is a materialized copy that only reflects the offline table when a publish runs. If the nightly job updates the offline table but no publish follows, the endpoint continues serving the older values. Republishing after the refresh aligns the online store with the current offline data and eliminates the stale predictions.

Why this answer

Feature Store keeps an offline table for training and an online table for serving. Updates to the offline table do not propagate automatically; a publish step materializes them into the online store. When the nightly refresh runs without a subsequent publish, the endpoint keeps reading the prior values, producing stale predictions for affected customers until the online store is refreshed.

Exam trap

The trap here is assuming the online feature store stays synchronized with the offline table without an explicit publish step.

81
MCQmedium

Which action should be taken to ensure that sensitive PII data is not leaked when logging models using MLflow?

A.Encrypt the entire MLflow experiment folder using workspace-level encryption.
B.Ensure data used for model artifacts is scrubbed of PII.
C.Set the MLflow model registry to 'Private' mode.
D.Delete the MLflow run immediately after deployment.
AnswerB

The most effective way to prevent leakage is to ensure that no PII is included in the training sets or metadata captured by MLflow. Data engineering pipelines should perform masking or removal before the data reaches the training notebook, effectively ensuring the model artifacts are compliant with privacy standards.

Why this answer

Sanitizing data before logging or excluding sensitive columns from the training set is a security best practice. MLflow logs metadata and often sample data. If PII is included, it is persisted in the artifact store.

By ensuring only non-sensitive features are logged, you comply with data privacy regulations and prevent unauthorized access to sensitive information within your model artifacts.

Exam trap

Candidates often suggest masking data at the serving layer or using encryption, forgetting that MLflow logs the training artifacts, so PII must be scrubbed before the logging process begins.

82
MCQhard

When deploying a model to a production endpoint, your team requires that the model is served from a specific, immutable version. How do you ensure this in MLflow?

A.Always point the endpoint to the 'Production' tag.
B.Specify the exact model version in the serving configuration.
C.Use the latest model artifact from the experiment folder.
D.Rebuild the model container on every request.
AnswerB

Hard-coding the version number in the serving config ensures the endpoint is immutable. Regardless of what happens to the 'Production' alias in the registry, the endpoint will continue to serve that specific version, which is the required approach for high-stakes production environments requiring strict version control.

Why this answer

Linking the deployment to a specific Model Registry version alias or version number is the only way to ensure immutability. When you serve a model, you should point to a specific version (e.g., version 5) rather than a dynamic alias like 'Production' if you need to guarantee that no subsequent transitions will change the behavior of the endpoint, thereby ensuring production stability.

Exam trap

Candidates often rely on dynamic tags like 'Production' alias, missing the requirement for strict immutability which demands referencing an exact version number.

83
MCQeasy

Which component of MLflow tracks the parameters, metrics, and tags associated with a specific training run?

A.MLflow Model Registry.
B.MLflow Tracking.
C.MLflow Projects.
D.MLflow Recipes.
AnswerB

MLflow Tracking provides the API and UI to log and visualize parameters, code versions, metrics, and output files. It acts as the central repository for experiment metadata, enabling teams to track and compare the performance of various model iterations effectively throughout the development cycle.

Why this answer

MLflow Tracking is the primary API used for logging and querying machine learning experiments. It allows data scientists to record key information during the training phase, which is vital for experiment reproducibility and comparing performance across different model iterations. Without MLflow Tracking, maintaining visibility into the history of experiments would be manual, fragmented, and prone to losing critical context.

Exam trap

Candidates often confuse MLflow Tracking with the MLflow Model Registry, incorrectly assuming that tracking is responsible for stage management or deployment rather than just recording experiment metadata.

84
MCQeasy

You need to ensure that a model deployed to a Databricks Model Serving endpoint can be rolled back quickly if it starts performing poorly. Which feature should you use?

A.Use MLflow Model Registry to transition the model version back to Staging, which automatically reverts the serving endpoint.
B.Delete the current model version from the registry and register the previous version as a new version.
C.Configure the endpoint to use a webhook that triggers a rollback job when performance metrics drop.
D.Enable multiple model versions on the endpoint and use traffic splitting to shift traffic back to a previous version.
AnswerD

Databricks Model Serving supports serving multiple model versions on a single endpoint and allows you to configure traffic splits. If the new version performs poorly, you can quickly shift traffic back to the previous version without redeploying, enabling fast rollback. This is the recommended approach for safe deployments and rollbacks.

Why this answer

Model Serving endpoints can host multiple model versions and use traffic splitting to route requests. If a new version underperforms, you can instantly shift traffic back to a previous version by adjusting the split, achieving a fast rollback without redeployment. Other methods are manual, slow, or do not directly control the serving endpoint.

Exam trap

The trap here is thinking that changing a model version's stage in the registry automatically updates the serving endpoint, but stage transitions and endpoint configurations are separate.

85
MCQeasy

A machine learning engineer needs to schedule a Databricks job that runs a Python wheel task to execute a packaged training pipeline on a recurring basis. The pipeline code is built into a wheel and stored in Unity Catalog volumes. Which job configuration correctly executes this workload?

A.Create a job with a notebook task that runs %pip install on the wheel and then calls the pipeline via subprocess from within the notebook.
B.Create a job with a Spark submit task that passes the wheel as a --py-files argument and runs the training script as the main class.
C.Create a job with a notebook task that contains a single cell calling dbutils.library.install on the wheel path, then imports the training module.
D.Create a job with a Python wheel task, specifying the wheel file location and the entry point as package_name.module_name, and attach a cluster or serverless compute.
AnswerD

A Python wheel task is designed exactly for this: you provide the wheel path and a fully qualified entry point such as package.module, and Databricks installs the wheel on the compute before invoking the entry point. This yields reproducible, versioned execution of the packaged training pipeline on a recurring schedule.

Why this answer

A Python wheel task is the purpose-built job task type for running packaged Python code: you point it at the wheel and a fully qualified entry point, and Databricks handles installation and invocation. Notebook-based installs or subprocess calls, and Spark submit tasks, do not provide the same reproducible, entry-point-driven execution.

Exam trap

The trap here is treating any method that installs a wheel as equivalent to a Python wheel task, when only the wheel task uses the packaged entry point directly.

86
MCQmedium

A data science team uses MLflow Tracking with a remote tracking server backed by a Databricks-hosted MySQL instance for the backend store and an Azure Data Lake Storage Gen2 path for artifacts. A model-training notebook writes metrics and a model artifact, then calls mlflow.register_model to promote the run into Unity Catalog. Reviewers report that the run's metrics appear in the experiment UI, but the model version in the registry cannot be loaded by the deployment job. The deployment job fails when it tries to download the artifact. Which action most directly resolves the deployment failure?

A.Re-run the training notebook with mlflow.set_tracking_uri pointing at a local file path so the artifact is written to the driver's disk.
B.Grant the deployment job's service principal READ privileges on the ADLS Gen2 artifact location, or configure a storage credential and external location that the deployment job can access.
C.Register the model again using the run ID instead of the artifact URI so the registry stores the artifact inline in the backend database.
D.Increase the --model-serve-timeout on the deployment cluster so the job has more time to fetch the artifact from the tracking server.
AnswerB

Artifacts are stored at the artifact location configured on the tracking server, not inside the backend database. When the deployment job lacks permission to that ADLS path, it cannot download the model, even though metadata and metrics are visible. Aligning the service principal's access or wiring a Unity Catalog storage credential and external location resolves the download failure.

Why this answer

MLflow separates metadata from artifacts: the backend store holds run and metric records, while the artifact store holds the model files. A deployment job that can read experiment metadata still needs its own credentials to the artifact location. Aligning the service principal's access, or exposing the path through Unity Catalog storage credentials and external locations, is what lets the artifact download succeed.

Exam trap

The trap here is assuming that because metrics and run metadata are visible in the UI, the artifacts are equally accessible to every consumer of the run.

87
Multi-Selecthard

A regulated financial services firm must prove that every model promoted to production on Databricks is traceable and governed. They use Unity Catalog for models and MLflow for experiment tracking. Which two practices most directly satisfy an auditor's requirement to trace a production model version back to its training data and code? (Choose two.)

Select 2 answers
A.Grant the production service principal MANAGE on the registered model so it can update versions as needed.
B.Enable the MLflow run's source notebook to be stored with the run so the exact code revision is retrievable from the run details.
C.Log the training dataset path or Delta table version and the Git commit SHA as tags and parameters on the MLflow run that produced the registered model version.
D.Increase the model version's stage description with a free-text summary of the business purpose and the approving manager's name.
E.Configure the serving endpoint to log every request and response payload to a Delta table for later inspection.
AnswersB, C

MLflow records the source notebook or Git reference with each run, which lets reviewers retrieve the code that generated the model. Combined with the run-to-version link in the registry, this closes the loop from production model back to the authored code, which is precisely what a traceability audit expects.

Why this answer

Traceability requires machine-readable links from the production model version back to the exact data and code. Recording the dataset version and Git commit on the source run, plus preserving the run's source notebook, gives auditors a verifiable chain. Descriptions, request logging, and permissions document or control access but do not establish provenance, so they cannot satisfy the requirement alone.

Exam trap

The trap here is treating documentation fields like version descriptions or access permissions as equivalent to recorded data and code lineage.

88
MCQhard

An ML engineer is instrumenting a Databricks job that trains and evaluates several models. They want each model's metrics, parameters, and artifacts grouped so that a downstream automated promotion step can compare candidates within the same experiment. Which MLflow practice best supports programmatic comparison across runs in this job?

A.Log all candidates as versions of a single registered model and compare them by reading each version's description field for metric values.
B.Write metrics to a Delta table outside MLflow and have the promotion step read that table, using MLflow only to store the model artifacts.
C.Create a separate experiment for each model candidate so that runs are isolated and cannot interfere with one another.
D.Log all candidates into a single experiment, tagging each run consistently so the promotion step can filter and rank runs by metric within that experiment.
AnswerD

Keeping candidates in one experiment lets the promotion step search runs by metric and filter with tags in a single query, producing a clean ranking of candidates. Consistent tags add the metadata needed to identify job, dataset, or model family, so automated comparison and selection are straightforward and auditable.

Why this answer

Grouping candidate runs in one experiment lets MLflow's search API rank and filter them by metric and tag in a single query, which is what an automated promotion step needs. Splitting experiments, overloading registry descriptions, or duplicating metrics elsewhere all break that integrated, queryable comparison.

Exam trap

The trap here is assuming that MLflow model versions carry metric data that can be ranked, when metrics live on runs within an experiment.

89
MCQhard

You are managing a Databricks environment and need to ensure that ML models are reproducible across different workspaces. Which strategy is most effective for cross-workspace model promotion?

A.Download the model artifact from the source workspace and upload it to the target workspace UI.
B.Use a central MLflow Model Registry in Unity Catalog to share model versions across workspaces.
C.Export the model as a pickled Python object and email it to the operations team.
D.Re-run the training notebook in each workspace to recreate the model artifact locally.
AnswerB

Unity Catalog's centralized model registry provides a single source of truth for all models, regardless of which workspace they are accessed from. This allows teams to promote models through environments (dev, staging, prod) while maintaining full auditability, lineage, and consistent governance over the model's lifecycle.

Why this answer

Using a centralized MLflow Model Registry via Unity Catalog allows models to be shared across workspaces securely. This avoids the manual export/import of artifacts and ensures that lineage, tags, and versioning remain consistent. By managing access through Unity Catalog permissions, you ensure that only authorized environments can read or promote specific models, creating a unified, compliant, and highly scalable model deployment lifecycle.

Exam trap

Candidates often suggest manual artifact copying or exporting/importing pickle files, failing to realize that Unity Catalog provides a centralized, secure, and native way to share models across workspaces.

90
MCQeasy

A data scientist wants to automate the retraining of a model whenever new data arrives in a Delta table. They need to orchestrate a multi-step workflow that includes data validation, feature engineering, model training, and deployment. Which Databricks feature should they use?

A.Databricks Jobs with a multi-task workflow
B.Databricks Repos
C.MLflow Projects
D.Delta Live Tables
AnswerA

Databricks Jobs support multi-task workflows, allowing you to define a directed acyclic graph of tasks with dependencies. You can schedule the job to trigger on a file arrival event or a schedule, and each task can run a notebook or Python script. This is the native orchestration tool in Databricks for building and automating ML pipelines, including retraining and deployment steps.

Why this answer

Databricks Jobs with multi-task workflows is the built-in orchestration service that allows you to define dependencies, schedule triggers, and monitor execution of complex ML pipelines. It can be triggered by file arrival events, making it ideal for retraining when new data lands. Other options like MLflow Projects, Delta Live Tables, and Repos serve different purposes and lack native orchestration capabilities.

Exam trap

The trap here is confusing code packaging or data pipeline tools with orchestration; only Databricks Jobs provides the scheduling and dependency management needed for end-to-end ML workflows.

91
MCQmedium

Your team uses MLflow Model Registry. A model version currently in Production has a critical flaw and must be rolled back to a previous version. The previous version is in the Archived stage. What is the most operationally sound approach to restore service quickly while preserving the audit trail?

A.Create a new registered model with the previous version's artifacts and point the serving endpoint to it.
B.Transition the flawed Production version to Archived, then transition the previous version from Archived to Production.
C.Delete the flawed Production model version and re-register the previous version as a new model version.
D.Leave the flawed version in Production and instead update the serving endpoint to load the previous version by its run ID.
AnswerB

Archiving the flawed version removes it from active serving while retaining its metadata and lineage, and moving the prior version back to Production restores the known-good artifact. This preserves a complete audit trail of stage transitions, which is essential for governance and post-incident review.

Why this answer

The correct approach is to archive the problematic Production version and promote the prior known-good version back to Production. This restores service using an already-validated artifact while maintaining full stage-transition history for audit and rollback traceability. Deleting versions, creating new models, or bypassing the registry all sacrifice governance or introduce unnecessary operational risk.

Exam trap

The trap here is assuming that a rollback requires creating a new model version or deleting the bad one, when stage transitions alone can restore service while preserving lineage.

92
MCQeasy

You are monitoring a model served on a Databricks Model Serving endpoint. You need to track the distribution of incoming request payloads to detect data drift. Which Databricks feature should you use to automatically capture and store inference logs for analysis?

A.Enable inference tables on the model serving endpoint.
B.Configure the endpoint to log to MLflow Tracking.
C.Use Databricks SQL dashboards to query the endpoint's access logs.
D.Enable model serving logs to be written to a cloud storage bucket.
AnswerA

Inference tables in Databricks Model Serving automatically capture the request payloads, response payloads, and metadata for each request to the endpoint. These logs are stored in a Delta table that you can query for monitoring and drift detection. This feature is designed specifically for capturing inference data without additional code.

Why this answer

Inference tables are a Databricks Model Serving feature that automatically logs request and response payloads to a Delta table. This enables easy querying and monitoring for data drift without writing additional code. Other options either do not capture payloads automatically or are not designed for this purpose.

Exam trap

The trap here is confusing MLflow Tracking with inference logging; MLflow Tracking is for training runs, not for capturing live serving requests.

93
Multi-Selecthard

A team is using Databricks Feature Store to manage features for a real-time fraud detection model. They need to ensure that the features used during training are consistent with those served at inference time. Which two actions should they take to achieve this? (Choose two.)

Select 2 answers
A.Publish the model with the Feature Store, so that the model automatically looks up the latest feature values from the online store at inference time.
B.Manually copy the feature values from the offline store to the online store before each inference request to ensure freshness.
C.Use the FeatureStoreClient to create a training set that joins features from the feature tables, ensuring point-in-time correctness.
D.Implement a custom UDF in the model to compute features on the fly from raw data, bypassing the Feature Store entirely.
E.Train the model using a separate notebook that reads directly from the online store to simulate inference-time feature retrieval.
AnswersA, C

When you log a model with Feature Store, it records the feature lookups. At inference time, the model uses the online store to fetch the latest feature values, ensuring that the same feature transformations are applied. This maintains consistency between training and serving and is the recommended practice for real-time models.

Why this answer

Using FeatureStoreClient.create_training_set ensures point-in-time correctness during training, and publishing the model with Feature Store ensures that inference uses the same feature lookups from the online store. Together, these actions guarantee that features are consistent between training and serving, which is critical for real-time fraud detection.

Exam trap

The trap here is thinking that manual synchronization or custom UDFs can replace the Feature Store's automated consistency mechanisms, when they actually introduce skew.

94
MCQeasy

Which component of Databricks is specifically designed to prevent training-serving skew by ensuring that feature engineering code is consistent during both model training and real-time inference?

A.MLflow Tracking Server
B.Databricks Feature Store
C.Unity Catalog
D.Databricks SQL
AnswerB

The Feature Store provides a unified API to compute and store features. When a model is logged with the Feature Store, the transformations are packaged with it, ensuring that identical code is applied to raw data at serving time, effectively eliminating training-serving skew in production pipelines.

Why this answer

The Databricks Feature Store acts as a centralized repository for features. By using the same feature definitions and transformation logic for both training and serving, it eliminates the discrepancy known as training-serving skew. This ensures that the features fed to the model at inference time are calculated exactly as they were during training, maintaining consistent performance and model reliability.

Exam trap

Candidates often confuse model registries with feature stores, failing to identify which specific component is responsible for eliminating training-serving skew.

95
MCQmedium

In an automated MLOps workflow, what is the best practice for handling model training failures?

A.Retrying the training job indefinitely.
B.Ignoring the error and deploying the previous model.
C.Logging the error and triggering an automated notification.
D.Deleting the failed experiment run.
AnswerC

This is the best practice. By capturing the error details and alerting the team, you ensure visibility and quick resolution. Automated logging preserves the context of the failure, which is crucial for debugging, while notifications ensure that the right people are aware of the production pipeline status.

Why this answer

Failing fast and sending alerts is the standard practice. Automated pipelines should include error handling that logs the failure cause to MLflow and triggers a notification via email or Slack. This ensures that data scientists can address the issue immediately without waiting for a manual audit of the pipeline, maintaining the overall health and reliability of the automated machine learning lifecycle.

Exam trap

Candidates often choose manual monitoring or retrying the training job immediately without logging, failing to recognize that automated pipelines require proactive error reporting and systemic logging for root cause analysis.

96
MCQmedium

You are retiring a real-time model endpoint on Databricks Model Serving. The endpoint has been serving production traffic for six months and you want to archive its request logs for compliance before deleting the endpoint. Which action should you take first?

A.Download the endpoint's model artifact from the MLflow Model Registry and store it as the audit record.
B.Use the Databricks REST API to list the endpoint's configuration and save the JSON response as the archive.
C.Enable inference table logging on the endpoint and wait for the next inference requests to be persisted.
D.Query the existing inference table that was configured for the endpoint and export its contents before deleting the endpoint.
AnswerD

When an endpoint is configured with an inference table, every request and response is written to a Delta table in Unity Catalog. That table persists independently of the endpoint, so querying and exporting it preserves the historical payloads needed for compliance before the endpoint itself is removed.

Why this answer

Model Serving writes inference payloads to a Delta inference table when the feature is enabled, and that table survives endpoint deletion because it is a separate Unity Catalog object. Exporting the table before removing the endpoint preserves the historical request and response records required for compliance, whereas configuration snapshots or model artifacts do not contain traffic data.

Exam trap

The trap here is assuming that enabling inference table logging at deletion time will backfill the historical requests that the endpoint already served.

97
MCQhard

Refer to the exhibit. The deployment pipeline is failing to load the model artifact in the target production environment. What is the most likely cause?

A.The model version is not set to 'PRODUCTION' in the metadata.
B.The path '/dbfs/...' is not a reliable way to reference artifacts across different clusters.
C.The conda_env file is missing from the artifact storage.
D.The run_id 'd7a8e9f0' has been archived and is no longer accessible.
AnswerB

Hard-coding DBFS paths is an anti-pattern. MLflow models should be referenced using the 'runs:/' or 'models:/' URI formats. These URIs are managed by the MLflow client and resolve correctly to the underlying artifact location, avoiding the path-resolution issues that occur when using literal file system paths.

Why this answer

The exhibit shows a hard-coded path beginning with '/dbfs/'. In Databricks, accessing DBFS via local file system paths can be inconsistent across different clusters or environments. The correct approach is to use the MLflow URI (e.g., 'runs:/...') to load models, as it abstracts the underlying storage location and ensures the artifact is resolved correctly regardless of the execution environment or storage configuration.

Exam trap

Many test-takers assume local or absolute DBFS paths are safe to use in deployment code, failing to recognize that MLflow URIs are required for environment portability.

98
MCQeasy

A team wants a scheduled Databricks job to automatically retrain a demand-forecasting model whenever the upstream feature table receives new data, and to register the resulting model version only if validation metrics improve. Which Databricks capability should they use to orchestrate this?

A.A cron expression on the training notebook that runs every minute to poll for new feature data.
B.A Delta Live Tables pipeline that materializes the feature table and automatically promotes models with the highest accuracy.
C.An MLflow webhook that fires when a new experiment run is created and calls the registration API.
D.A Databricks job with a table-update trigger that runs the training notebook and a conditional registration task.
AnswerD

Databricks jobs support triggers based on upstream table updates, which is exactly the event described. Chaining a training task with a task that conditionally registers the model based on metrics implements the gated promotion. This uses native orchestration, so no external scheduler or custom polling code is required.

Why this answer

Databricks jobs support table-update triggers, letting a run start when the feature table changes. A multi-task job can then train the model and use a conditional task to register the version only when validation metrics beat the current production baseline, all within native orchestration.

Exam trap

The trap here is reaching for MLflow webhooks or Delta Live Tables for event-driven retraining, when the table-update trigger on a Databricks job is the native mechanism.

99
MCQmedium

A data scientist has registered a model in Unity Catalog under the name `prod.ml.forecast_model`. They now need to define a service-level objective (SLO) that automatically monitors the model's prediction quality in production, not just endpoint uptime. Which Databricks feature should they configure to detect model performance degradation?

A.MLflow Model Registry webhooks
B.Cluster autoscaling policies on the serving cluster
C.Model Serving endpoint health checks
D.Lakehouse Monitoring on the inference table
AnswerD

Lakehouse Monitoring is Databricks' native solution for tracking data and model quality drift. By creating a monitor on the Delta inference table that stores model inputs and predictions, you get automatic profile and drift metrics, including model quality metrics when ground truth is joined. This directly satisfies the requirement to monitor prediction quality, not just infrastructure health.

Why this answer

Lakehouse Monitoring on the inference table is the only option that provides statistical monitoring of model inputs and outputs. It computes drift and profile metrics and, when ground truth labels are available, model quality metrics such as accuracy or RMSE. This makes it the correct mechanism for an SLO tied to prediction performance rather than simple endpoint availability.

Exam trap

The trap here is assuming that serving endpoint health checks or autoscaling provide model quality monitoring, when they only cover infrastructure availability and capacity.

100
MCQeasy

Which feature in Databricks allows you to automatically track training code, parameters, metrics, and models during the development phase?

A.Databricks Delta Lake
B.MLflow Tracking
C.Databricks Feature Store
D.Databricks Workflows
AnswerB

MLflow Tracking provides the API and UI to log training parameters, metrics, and code versions. It is specifically designed to manage the experimental phase of machine learning, allowing data scientists to organize runs, compare performance metrics, and keep track of model artifacts for later deployment in the Model Registry.

Why this answer

MLflow Tracking is the core component of Databricks for experiment management. It provides a structured way to log and visualize the results of machine learning runs. Understanding this component is fundamental to MLOps because it allows teams to compare models, track performance over time, and ensure that every experiment is reproducible, which is essential for professional model development and auditing.

Exam trap

Candidates frequently confuse MLflow Tracking with the MLflow Model Registry. Tracking is for experiment metadata and metrics during development, while the Registry is for lifecycle management and versioning.

101
MCQmedium

A data science team trains a scikit-learn model on a Databricks cluster and needs the same feature-engineering logic to run identically in a nightly batch scoring job and in a real-time Model Serving endpoint. They want a single artifact that encapsulates preprocessing and the estimator. Which approach should they use?

A.Log the model with MLflow using the sklearn flavor and a custom pyfunc wrapper that includes the preprocessing steps.
B.Register the raw estimator in Unity Catalog and reimplement preprocessing separately in the batch job and endpoint code.
C.Package the preprocessing into a custom container image and reference the image from the model version metadata.
D.Log the model with the sklearn flavor and pass the preprocessing function name as a signature parameter.
AnswerA

The pyfunc flavor wraps arbitrary Python logic, so preprocessing and the estimator travel as one artifact. Batch jobs load it with mlflow.pyfunc.load_model and the serving endpoint loads the same artifact, guaranteeing identical transformations. This is the standard Databricks pattern for eliminating training-serving skew when feature logic must be shared.

Why this answer

Wrapping preprocessing and the estimator in a single pyfunc model lets MLflow serialize both as one artifact. Batch scoring loads that artifact with the pyfunc loader, and Model Serving uses the same artifact, so feature transformations are guaranteed identical and training-serving skew is avoided without duplicating logic.

Exam trap

The trap here is believing a model signature or container image reference can carry executable preprocessing, when only a pyfunc wrapper actually bundles that logic into the artifact.

102
MCQmedium

Which TWO statements regarding the use of Unity Catalog in Databricks for MLOps are correct?

A.Unity Catalog enforces access control at the table level only, not on MLflow models.
B.It enables centralized lineage tracking between data tables and MLflow models.
C.Unity Catalog is only compatible with Python and does not support SQL-based model management.
D.Models registered in Unity Catalog can be shared across multiple workspaces.
E.Users must manually export model artifacts to an external store to share them via Unity Catalog.
AnswerB, D

Unity Catalog provides a unified lineage graph that shows how data flows from source tables into ML models. This visibility is essential for understanding the impact of data changes on model performance and for fulfilling regulatory requirements regarding data provenance and model transparency.

Why this answer

Unity Catalog acts as the centralized governance layer for all data and AI assets. It provides fine-grained access control and lineage tracking, allowing organizations to maintain visibility over which data is used by which models. By centralizing these assets, teams can collaborate safely while ensuring that compliance and auditing requirements are satisfied across the entire organization, regardless of the individual workspace where the work is performed.

Exam trap

Candidates often assume that models registered in Unity Catalog are isolated to a single workspace, missing that cross-workspace sharing and centralized governance are core advantages of using Unity Catalog for enterprise ML.

103
MCQmedium

What is the primary advantage of using a Model-as-Code approach in Databricks for machine learning deployments?

A.It ensures that the model is always trained on the latest data.
B.It enables consistent, reproducible, and versioned deployment environments.
C.It automatically generates unit tests for all machine learning code.
D.It eliminates the need for any monitoring of model performance.
AnswerB

By codifying everything from environment settings to deployment logic, teams ensure that the production environment is identical to the testing environment. This consistency is the foundation of reliable MLOps, enabling teams to deploy with confidence and revert to previous known-good states if issues arise in production.

Why this answer

Model-as-Code treats the entire deployment process—infrastructure, environment configuration, and code—as versioned artifacts. This allows for total reproducibility, where any production state can be rolled back or recreated using stored definitions. This practice is essential for enterprise MLOps, as it ensures that deployments are predictable, scalable, and audit-compliant, significantly reducing the risks associated with manual configuration changes in a production environment.

Exam trap

Candidates often describe Model-as-Code as merely 'automating deployments,' failing to emphasize the critical aspects of reproducibility, environment parity, and versioning that define the approach.

104
MCQhard

A machine learning team is using MLflow Model Registry to manage a model that is deployed to a production endpoint. They need to implement a CI/CD pipeline that automatically transitions a model version from 'Staging' to 'Production' only after it passes a set of validation tests. Which MLflow feature allows them to trigger the transition based on test results?

A.MLflow Model Registry stage transitions via the REST API
B.MLflow Projects with conditional steps
C.Databricks Jobs with task dependencies
D.MLflow Model Registry webhooks
AnswerA

The MLflow Model Registry REST API provides endpoints to transition model versions between stages. A CI/CD pipeline can call these endpoints programmatically after validation tests pass. This allows automated promotion based on test outcomes. The API is the standard way to integrate stage transitions into external workflows.

Why this answer

The MLflow Model Registry REST API allows programmatic stage transitions. In a CI/CD pipeline, after validation tests pass, you can call the API to move the model version from Staging to Production. This integrates seamlessly with automation tools.

Webhooks are for notifications, not for initiating transitions, and other options lack the direct capability to change stages based on test results.

Exam trap

The trap here is confusing webhooks, which react to transitions, with the API that actually performs the transition on demand.

105
MCQmedium

You are responsible for a fraud detection model deployed to a Databricks Model Serving endpoint. The model was trained on transaction data from the past 6 months. After two months in production, you notice a gradual decline in precision and recall. You suspect data drift. Which approach should you use to monitor feature drift for this endpoint?

A.Set up a Databricks SQL dashboard that queries the endpoint's logs and calculates the average prediction value over time.
B.Configure the endpoint to automatically retrain the model when drift is detected using the built-in auto-retrain feature.
C.Use MLflow Tracking to log the model's predictions and manually compare them to ground truth labels in a notebook.
D.Enable inference logging on the endpoint and use Databricks Lakehouse Monitoring to compare production feature distributions against the training baseline.
AnswerD

Inference logging captures the feature values sent to the endpoint, and Lakehouse Monitoring can compute drift metrics by comparing these to a baseline profile from training data. This provides automated drift detection and alerts, directly addressing the scenario.

Why this answer

Inference logging captures the actual feature values sent to the model endpoint, and Lakehouse Monitoring provides built-in drift metrics by comparing these to a baseline profile. This combination enables automated, scalable drift detection without manual intervention. Other options either misrepresent platform features or fail to statistically compare distributions.

Exam trap

The trap here is assuming that Model Serving includes automatic retraining or that simple prediction monitoring suffices for drift detection.

106
MCQmedium

You are implementing model monitoring to detect data drift. Which Databricks feature should you use to automatically track and alert on changes in the distribution of input data features over time?

A.MLflow Experiment Tracking
B.Delta Live Tables Expectations
C.Databricks Lakehouse Monitoring
D.Databricks SQL Alerts
AnswerC

Lakehouse Monitoring is the purpose-built service for detecting drift in input features and model predictions. It automatically creates dashboards and alerts by comparing production data against training baselines, providing essential visibility into model health after deployment in a production environment.

Why this answer

Databricks Lakehouse Monitoring provides automated insights into data quality and drift. By tracking statistical properties of incoming data against a baseline, it generates alerts when significant shifts are detected. This is a crucial component of MLOps, as it proactively identifies when a model's performance might degrade due to environmental changes, allowing teams to retrain or adjust models before they negatively impact business outcomes.

107
MCQmedium

Your organization requires that all models deployed to production must be signed by a security officer. How can you enforce this requirement within the Databricks MLflow Model Registry?

A.Delete all models in the Production stage and only re-upload the signed models.
B.Grant 'CAN_MANAGE' permissions on the registry to everyone to ensure transparency.
C.Restrict 'CAN_MANAGE' permissions on the model, and use an automated CI/CD pipeline for promotion.
D.Set the model version description to 'Signed by Security' after manual inspection.
AnswerC

Restricting permissions ensures that only authorized service principals or security officers can manage transitions. By requiring an automated CI/CD pipeline, you ensure that the promotion process is audited and validated against security policies before the model is moved to the Production stage, effectively enforcing the organization's requirements.

Why this answer

Enforcing approval workflows is crucial for compliance and risk management in MLOps. By using MLflow's permissions and stages, you can restrict who can promote models to 'Production'. Combining this with programmatic checks or external CI/CD gates ensures that a model cannot be deployed without the necessary authorization, preventing unauthorized or unvetted code from reaching production inference services.

Exam trap

Candidates often suggest using manual UI approvals, ignoring the requirement for automated, audit-compliant CI/CD pipelines that enforce security officer sign-offs through programmatic controls.

108
MCQmedium

A team uses MLflow Projects to package training code and runs jobs on Databricks clusters. They want to ensure that a job run today can be reproduced six months later with the same library versions, even if the cluster's base image and PyPI packages have changed. Which practice best achieves this?

A.Pin exact library versions in the project's conda.yaml or requirements.txt, and log the resulting environment specification with the run.
B.Rely on the Databricks Runtime version of the cluster, since Databricks Runtime images are immutable and never receive library updates.
C.Store the trained model artifacts in DBFS and re-run the training notebook from the workspace revision history.
D.Use the latest versions of all libraries at run time and rely on MLflow autologging to record which versions were used.
AnswerA

Pinning exact versions in the project environment file ensures the dependency resolver installs the same libraries at run time, and logging the environment specification with the run captures what was actually used. This combination makes the run reproducible months later regardless of changes to the base image or PyPI, because the pipeline can rebuild the same environment from the recorded specification rather than relying on floating versions.

Why this answer

Reproducibility across time requires controlling the environment, not just observing it. Pinning exact versions in the project's environment file guarantees the same dependencies are installed, and logging the environment specification with the run records the actual resolved environment for future rebuilding. Relying on runtime versions, notebook revisions, or autologging alone does not prevent dependency drift, so those approaches cannot ensure the same result six months later.

Exam trap

The trap here is confusing recording the environment with controlling it, since autologging captures versions but does not pin them for future runs.

109
MCQmedium

A machine learning engineer is using Databricks Model Serving to deploy a model that requires a custom Python library not available in the default environment. The engineer wants to ensure the endpoint uses the exact library version and that the deployment is reproducible. Which approach should the engineer take?

A.Install the custom library on all driver and worker nodes of the cluster used for model training, then deploy the model to Model Serving.
B.Use the Databricks REST API to upload the library to the Model Serving endpoint after deployment, then restart the endpoint.
C.Package the custom library as a Python wheel and include it in the model's conda environment or requirements file when logging the model with MLflow.
D.Create an init script that installs the custom library on the Model Serving cluster, and attach it to the endpoint configuration.
AnswerC

MLflow models can capture dependencies via a conda environment or requirements file. By packaging the custom library as a wheel and including it in the model's environment specification, Databricks Model Serving will install that exact version when deploying the model. This ensures reproducibility and availability of the custom library at inference time.

Why this answer

To ensure a custom library is available and reproducible in Databricks Model Serving, it must be included as a dependency when logging the MLflow model. Packaging the library as a wheel and adding it to the conda environment or requirements file ensures that Model Serving installs the exact version during deployment. This is the supported and recommended practice.

Exam trap

The trap here is assuming that libraries installed on a training cluster or via init scripts will carry over to the serverless Model Serving environment, which is not the case.

110
MCQmedium

Refer to the exhibit. Your automated CI/CD pipeline triggered a model deployment to production, but the job failed with the error shown. What is the most likely cause?

A.The model was registered, but the cluster lacks permissions to read it.
B.The deployment job is referencing a model version that hasn't been created yet.
C.The model serving endpoint is already occupied by another model version.
D.The training cluster ran out of memory during the model artifact upload.
AnswerB

The REST exception confirms the system attempted to fetch a non-existent entity. In CI/CD, this suggests a dependency failure where the deployment script executes before the training job successfully completes the registration, indicating a need for improved job orchestration or explicit completion gating.

Why this answer

The error indicates that the deployment pipeline is attempting to access a model version that does not exist in the Model Registry. This usually happens when the pipeline triggers before the registration process completes or when a race condition occurs between the training job and the deployment job. Validating the existence of the model version before initiating deployment is essential for pipeline reliability.

Exam trap

Candidates often blame infrastructure or network issues. However, in automated CI/CD, the most common failure is a race condition where the deployment script executes before the registration job finishes.

111
MCQhard

A fraud detection model is served on a Databricks Model Serving endpoint. You notice that predictions for the same input vector differ between two consecutive requests within seconds, and there is no feature store or external cache involved. The model was logged with a fixed random seed and deterministic inference code. Which action should you take first to diagnose the inconsistency?

A.Convert the model to a different flavor such as ONNX and re-register it, because the original flavor cannot guarantee deterministic inference.
B.Increase the endpoint's concurrency and enable autoscaling, since differing predictions indicate the replicas are running different code.
C.Inspect the endpoint's request logs to confirm the payloads are identical and check whether the model version serving traffic changed between requests.
D.Retrain the model with a different random seed and redeploy, because seed instability is the likely cause of varying predictions.
AnswerC

Non-determinism in a supposedly deterministic model most often comes from the serving layer, not the model math. Verifying that the request payloads are byte-identical rules out client-side variation, and checking whether the endpoint switched model versions or routed to a different version reveals the most common cause of divergent outputs. This is the cheapest, highest-signal first step before deeper debugging of the model artifact or environment.

Why this answer

When a deterministic model returns different outputs for the same input, the fastest path to the cause is to verify the two requests were truly identical and to confirm which model version handled each. A version transition during a rollout, or a payload difference such as field ordering or missing values, explains the symptom without any model retraining. Retraining, scaling, or changing flavors are premature and do not produce diagnostic evidence.

Exam trap

The trap here is jumping to model-level explanations like seeds or flavors, when serving-layer causes such as version routing or payload differences are far more likely.

112
MCQhard

When auditing an ML pipeline in Databricks for compliance and governance, which THREE of the following should be verified?

A.The Git commit hash associated with the specific training run.
B.The personal credentials of the data scientist who manually ran the training.
C.The lineage of the training data, including source table versions.
D.The list of authorized users who can transition models to 'Production'.
E.The raw text of every email sent between the data scientists.
AnswerA, C, D

Traceability starts with code. By linking every model to a specific Git commit, auditors can verify the exact code state that produced the model. This is fundamental for reproducibility and ensuring that no unauthorized or unreviewed changes were introduced into the model generation process.

Why this answer

Auditing requires proof of data lineage, code versioning, and access control. Verifying these components ensures that the organization can prove how a model was built, what data was used, who approved it, and that the code was properly reviewed. This level of traceability is non-negotiable for regulated industries and essential for enterprise-grade MLOps maturity.

Exam trap

Candidates often include irrelevant metrics like model accuracy or latency as audit requirements, failing to realize that compliance audits focus on traceability, lineage, and access control.

113
MCQhard

Your organization is implementing an MLOps strategy that requires strict model governance. You need to ensure that no model is deployed to production unless it has been tagged with 'validated=true' in the MLflow Model Registry. How can you enforce this policy within your CI/CD workflow?

A.Use Databricks workspace permissions to prevent unauthorized users from deploying models.
B.Add a validation script in the CI/CD pipeline that queries the model version tags.
C.Rely on the MLflow UI to visually verify the tag before clicking 'Deploy'.
D.Set the default stage of all new models to 'Production' via a global config.
AnswerB

A custom validation script acts as a gatekeeper. By using 'mlflow.tracking.MlflowClient()' to retrieve the model version and inspect its dictionary of tags, the pipeline can verify the 'validated=true' condition. If the condition is not met, the pipeline fails, preventing the deployment from occurring.

Why this answer

Implementing a policy-based deployment gate in your CI/CD pipeline ensures that governance requirements are met before code or models hit production. By querying the model's tags using the MLflow client before attempting a deployment, you create a hard stop for non-compliant models. This is a critical security and compliance practice, preventing unauthorized or untested models from being served to end-users or critical business processes.

114
MCQmedium

Which of the following describes the purpose of a 'Gold' table in the Medallion architecture within an MLOps pipeline?

A.It stores raw data ingested directly from external sources.
B.It contains validated, business-level data ready for ML model training.
C.It serves as a transient staging area for schema evolution.
D.It stores model artifacts and hyperparameters for versioning.
AnswerB

The Gold layer serves as the final consumption layer, providing reliable, high-quality data. In an MLOps context, this is the optimal source for training data, ensuring that the model is built on clean and consistent information, which minimizes the risk of garbage-in, garbage-out performance issues.

Why this answer

The Gold layer contains highly processed, business-level data that is ready for consumption by downstream ML models or analytical applications. By ensuring data is clean, aggregated, and validated at this stage, MLOps teams can rely on high-quality features for training, which directly improves model performance and reduces the complexity of the feature engineering step in the training pipeline.

Exam trap

Candidates confuse Gold tables with raw ingestion tables (Bronze) or intermediate cleaned tables (Silver), missing that Gold stores analytics-ready business data.

115
MCQeasy

A data scientist wants to compare multiple runs within the same MLflow experiment to identify the best-performing model based on a custom metric. What is the most efficient way to do this in the MLflow UI?

A.Export all runs to a CSV file using the MLflow CLI and then analyze the CSV in a separate tool.
B.Use the MLflow UI's experiment page to sort runs by the custom metric and compare them side by side.
C.Query the MLflow tracking server's REST API and write a custom script to parse and rank the runs.
D.Open each run individually and manually record the custom metric values in a spreadsheet.
AnswerB

The MLflow UI experiment page lists all runs and allows sorting by any logged metric, including custom ones. You can select multiple runs and compare their parameters, metrics, and artifacts side by side. This is the intended and most efficient way to identify the best run.

Why this answer

The MLflow UI experiment page is designed for comparing runs: it displays all runs in a table, supports sorting by any metric, and allows selecting multiple runs for side-by-side comparison of parameters, metrics, and artifacts. Manual recording, CSV export, or custom API scripts are less efficient and unnecessary for this common task.

Exam trap

The trap here is overlooking the built-in comparison features of the MLflow UI and instead reaching for manual or programmatic methods that add unnecessary work.

116
Multi-Selecthard

A team is using Databricks Feature Store to manage features for a real-time model served via Databricks Model Serving. They need to ensure that the online feature values used at inference time are consistent with the training data. Which TWO practices should they implement? (Choose two.)

Select 2 answers
A.Log the model with the feature store's training set specification to enable automatic feature lookup at serving time.
B.Store all features in a single Delta table and query it directly from the model serving endpoint.
C.Use point-in-time lookups when creating the training dataset to avoid leaking future data.
D.Publish feature tables to an online store that is updated with the same pipeline that writes to the offline store.
E.Manually copy feature values from the offline store to the online store on a daily basis using a notebook.
AnswersC, D

Point-in-time lookups ensure that training examples only use feature values available at the time of the label event, mimicking real-time inference conditions. This prevents data leakage and aligns training with serving. Databricks Feature Store provides time-series lookups for this purpose, which is essential for temporal consistency.

Why this answer

Consistency between online and offline features is achieved by using the same pipeline to publish to both stores and by using point-in-time lookups during training. These practices ensure that the transformations and temporal alignment match between training and inference. Manual copying or direct Delta queries do not provide the required consistency or performance.

Exam trap

The trap here is thinking that logging the model with the training set specification alone guarantees consistency, when the online store must also be updated with the same pipeline.

117
MCQmedium

A data scientist at your company has trained a scikit-learn model and logged it with MLflow using the default 'sklearn' flavor. The model must now be deployed to a Databricks Model Serving endpoint for real-time inference, and the endpoint must automatically reload the new version whenever the model's stage changes to 'Production' in the MLflow Model Registry. Which deployment approach should the team use?

A.Log the model to a new MLflow experiment and configure the endpoint to poll the experiment's artifact location for changes.
B.Package the model into a Docker image with a Flask app and deploy it to a Databricks cluster running as a job.
C.Export the model artifact to DBFS as a pickle file and create a Python script that loads the pickle at endpoint startup.
D.Register the model version in the MLflow Model Registry, then create the serving endpoint by specifying the model name and the 'Production' stage in the endpoint configuration.
AnswerD

Databricks Model Serving natively integrates with the MLflow Model Registry. When you configure an endpoint with a registered model name and a stage such as 'Production', the endpoint automatically serves the version currently in that stage and reloads when the stage transitions, satisfying the automatic-reload requirement without custom code.

Why this answer

The requirement is automatic reload when a model version enters the 'Production' stage. Databricks Model Serving supports this directly by referencing the registered model name and stage in the endpoint configuration. Other approaches either bypass the registry, require manual redeployment, or rely on mechanisms the platform does not provide, making them unsuitable for a managed real-time endpoint with stage-driven updates.

Exam trap

The trap here is assuming the endpoint must be recreated or manually updated each time a model version changes stage, when stage-based endpoint configuration already reloads automatically.

118
MCQmedium

You are building a Databricks job that trains a model and registers it to the MLflow Model Registry. After registration, you need to automatically transition the model version to 'Staging' only if its validation accuracy exceeds 0.95. The job should fail if the accuracy is below this threshold. Which approach should you implement?

A.Use the MLflow Client to register the model, then use a conditional step in the job to call the transition_model_version_stage method if the accuracy metric meets the threshold.
B.Configure the MLflow Model Registry webhook to automatically transition the model version to 'Staging' when a new version is registered.
C.Use the Databricks REST API to update the model version stage after the job completes, relying on the job's success or failure to indicate accuracy.
D.Set the model version's stage to 'Staging' during registration using the stage parameter in mlflow.register_model, and rely on the training script to skip registration if accuracy is low.
AnswerA

The MLflow Client provides programmatic control over model versions, including transition_model_version_stage. By retrieving the run's metrics and conditionally transitioning only when accuracy exceeds 0.95, you enforce the business rule. If the condition fails, you can raise an exception to fail the job. This approach is flexible and integrates well with Databricks jobs.

Why this answer

To conditionally transition a model version based on a metric, you need to programmatically retrieve the metric and then call the MLflow Client's transition method. This allows you to enforce a threshold and fail the job if the condition is not met. Webhooks and REST API calls lack the ability to evaluate metrics before transitioning, and setting the stage during registration does not incorporate a conditional check.

Exam trap

The trap here is assuming that webhooks or registration-time stage settings can enforce a metric threshold, when they actually operate after or independently of metric evaluation.

119
MCQhard

Which THREE of the following are primary responsibilities of an MLOps engineer when maintaining production ML models in Databricks?

A.Monitoring model performance and detecting data drift.
B.Writing the core machine learning research papers for the organization.
C.Establishing CI/CD pipelines to automate testing and deployment.
D.Configuring workspace security and access control for model artifacts.
E.Manually retraining every model daily regardless of performance metrics.
AnswerA, C, D

Drift detection is critical to maintaining model accuracy. As data distributions change, models can lose predictive power. MLOps engineers must implement monitoring solutions to identify these shifts early, allowing for timely retraining and redeployment, which protects the organization from deploying degraded models to production users.

Why this answer

An MLOps engineer's role is to ensure stability, performance, and compliance. This includes monitoring model performance to detect drift, managing the CI/CD pipeline for automated deployments, and ensuring security via robust access controls. By balancing these tasks, they ensure that the machine learning system remains reliable and valuable to the business over time, effectively bridging the gap between development and operations.

Exam trap

Candidates focus exclusively on model training metrics, neglecting essential MLOps operational tasks like drift monitoring, CI/CD, and security controls.

120
MCQmedium

Your team is using Databricks Feature Store to serve features to a real-time model deployed on Databricks Model Serving. A data scientist updates the feature computation logic for one of the features and publishes a new version to the online store. However, the model endpoint continues to return stale feature values for that feature. What is the most likely cause?

A.The model endpoint has not been restarted after the feature table update, so it is still using cached feature values.
B.The online store is not configured to automatically sync with the offline store, so the new feature values were never published.
C.The model was logged with a feature spec that references an older version of the feature table, so it continues to use the old feature values.
D.The model was logged without the feature lookup package, so it cannot query the online store.
AnswerC

When you log a model with Feature Store, the feature spec captures the exact feature table versions at that time. If you later publish a new version of a feature table, the existing model still references the version recorded in its feature spec. To use the updated feature, you must re-log the model with a new feature spec that points to the new version, then redeploy the endpoint.

Why this answer

The model's feature spec records the versions of the feature tables used during training and logging. When a new feature table version is published, existing models still reference the old version, leading to stale values at inference. To resolve this, the model must be re-logged with an updated feature spec that includes the new version, and the endpoint must be updated with the new model version.

Exam trap

The trap here is assuming that updating the feature table automatically updates all models that use it, when in fact each model's feature spec pins specific table versions.

121
Multi-Selecthard

A machine learning engineer is responsible for monitoring a production model deployed to Databricks Model Serving. The model predicts customer churn and is served via a REST endpoint. The engineer needs to detect data drift and model performance degradation over time. Which TWO actions should the engineer take to enable effective monitoring? (Choose two.)

Select 2 answers
A.Enable automatic model retraining triggered by any change in the input data schema.
B.Enable inference logging on the model serving endpoint to capture request and response payloads.
C.Use the model's training accuracy as a proxy for production performance and set up alerts based on that metric.
D.Configure the endpoint to use a smaller instance type to reduce cost, as monitoring does not require additional resources.
E.Schedule a Databricks job to periodically compute drift metrics by comparing logged inference data with the training dataset.
AnswersB, E

Inference logging captures the input features and model predictions for each request, which is essential for computing drift metrics and evaluating performance over time. Databricks Model Serving allows you to enable inference logging to a Delta table, where the data can be analyzed using Databricks SQL or notebooks. Without this logging, there is no historical record of production data to compare against training data or to compute accuracy metrics, making drift detection impossible.

Why this answer

Effective monitoring of a production model requires capturing inference data and analyzing it for drift and performance. Enabling inference logging provides the raw data, while scheduling a job to compute drift metrics against the training set enables detection of data drift. Together, these actions allow the engineer to identify when the model's input distribution diverges from training, which is a key signal of degradation.

Exam trap

The trap here is thinking that training accuracy or automatic retraining can substitute for actual production monitoring, when in fact you must capture and analyze live inference data to detect drift and degradation.

122
MCQmedium

Your ML pipeline requires a complex environment with specific C++ dependencies. What is the recommended way to manage this in Databricks?

A.Install dependencies in a setup script at runtime.
B.Use Databricks Container Services with a custom image.
C.Include dependencies in the MLflow model artifact.
D.Create a shared library JAR file.
AnswerB

Custom Docker images allow you to pre-install complex C++ dependencies, system libraries, and specific compilers that aren't available in the standard Databricks runtime. This ensures that every cluster node starts with an identical, fully configured environment, providing the stability and reproducibility required for complex MLOps pipelines.

Why this answer

Databricks Container Services allows users to define custom Docker images as the execution environment. This is the only way to handle complex system-level dependencies like C++ libraries that cannot be installed via standard pip or conda commands. By using custom containers, teams ensure that the training and inference environments are identical and include all necessary system-level components.

Exam trap

Candidates often choose standard cluster libraries or init scripts, forgetting that complex system-level C++ dependencies require custom Docker images via Databricks Container Services.

123
MCQeasy

Which component of Databricks ML is best suited for tracking model hyperparameters, metrics, and code versions during the experimentation phase of the ML lifecycle?

A.Databricks SQL
B.MLflow Tracking
C.Delta Lake
D.Unity Catalog
AnswerB

MLflow Tracking provides the API and UI to record experiments. It allows developers to log parameters, performance metrics, and artifacts like serialized models and plots, which are essential for comparing runs and selecting the optimal model for production deployment.

Why this answer

MLflow Tracking is the dedicated component designed to record every experiment's parameters, metrics, and code state. By capturing these elements, data scientists can compare different iterations of a model to identify the best-performing configuration. This visibility is vital in MLOps, as it creates an audit trail of how a final model was derived, supporting reproducibility and informed decision-making before promoting a model to the registry.

124
MCQhard

You are using Databricks Feature Store to manage features for a model. A feature table is updated daily with new data. Your model training job reads from the feature table and logs the model with MLflow. To ensure that the model in production uses the correct feature values at inference time, what must you do when logging the model?

A.Log the model with a static copy of the feature values used during training.
B.Log the model with the feature table's primary key and timestamp columns only.
C.Log the model with the feature store lookup tables and specify the feature table names.
D.Log the model without any feature store metadata and manually join the feature table in the scoring job.
AnswerC

When logging a model that uses Databricks Feature Store, you must include the feature store metadata by specifying the feature table names. This allows the model to automatically look up the latest feature values at inference time. MLflow integrates with Feature Store to package this information, ensuring consistency between training and serving.

Why this answer

To ensure the model uses the correct feature values at inference time, you must log the model with the feature store lookup tables and specify the feature table names. This enables automatic feature retrieval, maintaining consistency between training and serving and ensuring that daily updates are reflected in predictions.

Exam trap

The trap here is assuming that logging the model alone is sufficient, or that static features are acceptable, when in fact Feature Store integration requires explicit metadata to enable automatic lookups.

125
MCQmedium

Your team uses MLflow Projects to package training code. A colleague runs the project on a Databricks cluster and it fails with a dependency conflict because the cluster has an older version of a library than the project's conda environment specifies. What is the most reliable way to ensure the project uses its declared dependencies without modifying the shared cluster?

A.Run the MLflow project with the --experiment-name flag to isolate dependencies per experiment.
B.Use the MLflow Projects CLI with the --conda flag or run via mlflow run, which creates an isolated conda environment from the project's conda.yaml.
C.Install the required library version globally on the cluster and restart it before running the project.
D.Convert the project to a Databricks notebook and use %pip install at the top of the notebook.
AnswerB

MLflow Projects supports conda environments defined in conda.yaml. Running with mlflow run and conda enabled creates an isolated environment that installs the declared dependencies, avoiding conflicts with the cluster's preinstalled libraries. This is the intended mechanism for reproducible dependency management.

Why this answer

MLflow Projects are designed to encapsulate dependencies in a conda.yaml, and running them with conda enabled creates an isolated environment that installs the specified versions. This avoids conflicts with shared cluster libraries and preserves reproducibility. Flags controlling experiment names, global installs, or notebook conversions do not provide the same isolation or reliability.

Exam trap

The trap here is assuming that any MLflow CLI flag or notebook-level pip install provides the same dependency isolation as the project's conda environment.

126
Multi-Selectmedium

You are designing a CI/CD pipeline for a machine learning model on Databricks. The pipeline must automatically retrain the model when new data arrives, validate it, and deploy it to a serving endpoint if it passes. Which two components are essential to achieve this? (Choose two.)

Select 2 answers
A.MLflow Tracking to log parameters, metrics, and artifacts for each run.
B.A Databricks Repos integration to version control notebooks and manage code changes.
C.A Databricks Model Serving endpoint that can be updated with the new model version via the MLflow API.
D.A feature store that automatically ingests new data and updates features in real time.
E.A Databricks Job that runs the training and validation notebooks on a schedule or trigger.
AnswersC, E

A Model Serving endpoint is essential for deploying the model and serving predictions. The pipeline must update the endpoint with the new model version after validation passes. This is typically done via the MLflow API or Databricks CLI. Without a serving endpoint, the model cannot be consumed in real time, so this is a critical component of the deployment step.

Why this answer

A Databricks Job provides the orchestration to run training and validation when new data arrives, and a Model Serving endpoint is required to deploy the validated model for real-time predictions. Together, they form the backbone of an automated retraining and deployment pipeline. MLflow Tracking, feature stores, and Repos are valuable but not strictly essential for the pipeline to function.

Exam trap

The trap here is overemphasizing supporting tools like MLflow Tracking or feature stores, which are helpful but not essential for the core automation and deployment loop.

127
Multi-Selecthard

You are auditing a Databricks environment to ensure compliance. Which TWO actions ensure the highest level of model lineage and reproducibility for models registered in MLflow?

Select 2 answers
A.Include the Git hash in the log_model metadata.
B.Use random seeds for all ML libraries.
C.Log the environment configuration (pip/conda) with the model.
D.Store all training data as CSV files in the workspace.
E.Manually copy artifacts to a shared folder.
AnswersA, C

Recording the Git hash during logging directly links the model artifact to the exact codebase used for training. This enables developers to checkout the precise state of the repository at the time of training, ensuring that lineage is preserved and debugging becomes straightforward.

Why this answer

Linking models to specific Git commits and tracking the environment dependencies via Conda or pip requirements files are essential for reproducibility. These steps ensure that when a model is redeployed, the exact code version and environment can be recreated, fulfilling the core MLOps requirement of auditability and consistent model behavior across different computing environments.

Exam trap

Candidates often focus on model accuracy metrics, ignoring that compliance and reproducibility require linking the model to specific code versions (Git) and exact dependency environments (Conda/pip).

128
MCQhard

Your organization requires a feature store strategy that supports low-latency point-in-time lookups for online inference. Which implementation approach best addresses this?

A.Using a temporary Spark SQL view for features.
B.Utilizing the Databricks Feature Store with online store integration.
C.Caching all feature data in local cluster memory.
D.Performing manual joins in inference notebooks.
AnswerB

The Databricks Feature Store is specifically designed to manage feature pipelines, ensuring data consistency between offline training and online serving. Integrating with a high-performance online store allows for low-latency lookups, while the built-in point-in-time join logic prevents training-serving skew and data leakage.

Why this answer

Databricks Feature Store enables point-in-time joins to avoid data leakage during training and supports publishing to low-latency online stores (e.g., DynamoDB or Cosmos DB). By using a centralized feature store, teams ensure consistency between training and serving. Point-in-time lookups are crucial to ensure that features used for predictions reflect exactly what was known at that specific moment, preventing bias.

Exam trap

Candidates often confuse batch feature tables with online store deployments, failing to recognize that low-latency online inference requires dedicated online store integration.

129
MCQmedium

When deploying a model as a real-time REST endpoint on Databricks, how can you ensure the infrastructure scales automatically to handle increased request traffic?

A.Configure the cluster to use a fixed number of workers in the cluster settings.
B.Deploy the model using Databricks Model Serving endpoints.
C.Manually add instances to the inference cluster using the Databricks API before peak times.
D.Enable 'Auto-terminate' on the inference cluster configuration.
AnswerB

Databricks Model Serving endpoints are designed specifically for high-performance, real-time inference. They automatically handle infrastructure provisioning and scaling, adjusting the number of active model instances based on the current load. This abstracts away the complexity of cluster management, ensuring reliable and efficient model serving in production.

Why this answer

Databricks Model Serving provides serverless, auto-scaling inference endpoints. By default, it manages the underlying infrastructure, scaling the number of replicas based on real-time request volume. This is a key MLOps requirement to ensure high availability and performance without manual intervention, allowing teams to handle unpredictable traffic spikes without over-provisioning and incurring unnecessary costs during low-traffic periods.

Exam trap

Candidates often choose manual cluster configuration or custom scaling scripts, failing to realize that Model Serving endpoints provide native, serverless auto-scaling without needing manual infrastructure overhead.

130
MCQmedium

A team runs a weekly retraining job that produces a new model version in Unity Catalog. Their production endpoint is currently serving version 4. They want to promote version 5 with zero downtime and the ability to roll back instantly if error rates rise. Which approach best meets these requirements?

A.Delete version 4 from the registry, deploy version 5, and recreate the endpoint so it picks up the new version automatically.
B.Configure the serving endpoint to serve version 5 and rely on the platform's rolling update so traffic shifts gradually while the previous configuration remains available for rollback.
C.Set the endpoint's served entities to version 5 with a traffic percentage of zero, then raise traffic to one hundred percent after validation.
D.Create a second endpoint for version 5, run both endpoints in parallel indefinitely, and split traffic manually at the load balancer.
AnswerB

Databricks Model Serving performs rolling updates that keep the endpoint available while new model versions warm up, and the prior configuration can be restored if problems appear. Pointing the endpoint at the new version and letting the rolling update proceed gives zero-downtime promotion with a fast rollback path, which matches both stated requirements without extra tooling.

Why this answer

Model Serving supports rolling updates that keep the endpoint live while a new model version loads, and the previous configuration remains restorable. Repointing the endpoint to the newer version and letting the rolling update run achieves zero-downtime promotion with a straightforward rollback. Deleting the old version or duplicating endpoints introduces risk or cost without improving the outcome.

Exam trap

The trap here is conflating a rolling update on a single endpoint with running duplicate endpoints or deleting the prior version.

← PreviousPage 2 of 2 · 130 questions total

Ready to test yourself?

Try a timed practice session using only ML Ops questions.

CCNA ML Ops Questions — Page 2 of 2 | Courseiva