Courseiva

CCNA Ml Workflows Questions

75 of 83 questions · Page 1/2 · Ml Workflows topic · Answers revealed

1
MCQeasy

A data scientist is using MLflow Tracking to log experiments. They want to compare multiple runs of a scikit-learn model and identify the run with the lowest RMSE. Which MLflow feature should they use?

A.The MLflow Tracking UI, which allows sorting and filtering runs by metrics.
B.The MLflow Projects feature, which packages code for reproducible runs.
C.The MLflow Model Registry, which stores model versions and their stages.
D.The MLflow Recipes (formerly Pipelines) framework, which automates model training.
AnswerA

The MLflow Tracking UI provides a table of runs with sortable and filterable columns, including metrics. You can sort by RMSE ascending to find the lowest value. It also supports parallel coordinates plots and metric charts for visual comparison. This is the primary tool for comparing runs in an experiment.

Why this answer

The MLflow Tracking UI is designed for visualizing and comparing runs within an experiment. It allows sorting by metrics such as RMSE, filtering, and generating charts. Projects, Model Registry, and Recipes serve different purposes: packaging, lifecycle management, and automation, respectively.

Therefore, the Tracking UI is the correct tool for this comparison.

Exam trap

The trap here is confusing MLflow components: Projects and Recipes are about packaging and automation, while the Tracking UI is for run comparison.

2
MCQmedium

You are tracking a deep learning model experiment using MLflow. You want to ensure that the model architecture and all hyperparameters are easily reproducible. Which approach is best practice?

A.Only log the final model artifacts to the MLflow Model Registry to save disk space.
B.Log parameters, metrics, and include the conda environment configuration with the model.
C.Manually copy the training code into a text field in the MLflow UI for every run.
D.Use a global variable in your notebook to store all hyperparameters for easy access.
AnswerB

Logging parameters and environment configuration ensures that the code can be executed in an identical environment. MLflow automatically packages these dependencies, allowing you to restore the training state precisely. This reproducibility is a fundamental requirement for compliance, debugging, and iterative model development in enterprise ML pipelines.

Why this answer

Logging model parameters and the environment configuration (like conda.yaml or requirements.txt) via MLflow ensures reproducibility. By capturing the exact package versions and architectural settings during the run, you eliminate the 'it works on my machine' problem. This practice is essential in production workflows where model drift or performance degradation requires auditing the exact conditions under which a specific model version was trained.

Exam trap

Candidates frequently log only metrics, ignoring the environment configuration. They fail to realize that without the exact package versions, a model cannot be reliably reproduced in a different environment.

3
MCQmedium

Refer to the exhibit. A data scientist receives this error while trying to load a model. What is the most likely cause of this failure in the workflow?

A.The model has been deleted from the registry.
B.The model version number is incorrectly referenced in the code.
C.The cluster does not have permission to access the MLflow registry.
D.The MLflow tracking URI is pointing to a different workspace.
AnswerB

The error explicitly states that the version does not exist. This is a common issue when using hard-coded version integers. If the model was never registered as version 5, or if it was deleted, the retrieval command will fail, confirming the code is pointing to a non-existent artifact.

Why this answer

This error indicates that the code is attempting to reference a specific version of a model that does not exist in the Model Registry. This often occurs when a script relies on hard-coded version numbers rather than dynamic lookups or stage-based references. In a production workflow, using stage aliases like 'Production' is preferred to ensure that the application always fetches the current valid version without needing manual updates when models are updated.

Exam trap

Candidates often assume the error is related to network connectivity or permissions, failing to notice that hard-coded version numbers are brittle and prone to breaking when versions are updated.

4
Multi-Selectmedium

Which THREE actions are essential when preparing a machine learning model for deployment using the Databricks Model Registry?

Select 3 answers
A.Transitioning the model version to the appropriate stage.
B.Logging the model with an input signature.
C.Adding metadata and descriptions to the model version.
D.Manually downloading the model pickle file to a local machine.
E.Hard-coding the model URI in the application inference code.
AnswersA, B, C

Transitions allow for the formal advancement of a model from experimentation to deployment. Using stages like 'Staging' and 'Production' enables teams to control which versions are used in downstream inference applications, ensuring only tested and approved artifacts are exposed to live data and production environments.

Why this answer

Preparing a model involves validating its performance, documenting its lineage, and assigning it to the appropriate registry stage. By logging the model with its signature, you ensure that the input/output schema is preserved. Transitioning through stages (Staging, Production) allows for controlled releases, while adding metadata via tags and descriptions provides necessary context for other stakeholders to understand the model's purpose, limitations, and performance characteristics in a production environment.

Exam trap

Candidates often overlook the importance of the input signature, incorrectly assuming that code alone is sufficient for deployment without defining the expected data schema for downstream inference.

5
MCQmedium

An ML engineer has a scikit-learn model trained locally and wants to log it to MLflow with a signature so that Databricks Model Serving can enforce input schema validation. The engineer calls mlflow.sklearn.log_model(model, 'model') but does not use infer_signature. What is the most accurate consequence when the model is later served on Databricks Model Serving?

A.The model will be served with a default signature that expects a single double column named 'features'.
B.Databricks Model Serving will automatically infer the signature from the first request payload and cache it for later validation.
C.Databricks Model Serving will reject the model version because a signature is mandatory for all Unity Catalog registered models.
D.The model will be served, but the endpoint cannot validate incoming request schemas against a stored signature, so malformed inputs may reach the model.
AnswerD

MLflow signatures are recorded only when log_model is called with a signature argument or infer_signature. Without it, the logged model has no stored input schema, so Databricks Model Serving cannot enforce request validation and will pass payloads directly to the model, where unexpected columns or types may cause runtime errors.

Why this answer

Logging a model without infer_signature leaves the MLflow model metadata without an input schema. Databricks Model Serving uses that stored signature to validate request payloads, so the endpoint will accept requests but cannot reject malformed ones. To gain schema enforcement, the engineer should call infer_signature on a sample DataFrame and pass it to log_model.

Exam trap

The trap here is assuming that Databricks Model Serving can infer or enforce a schema from live traffic even when no signature was logged with the model.

6
MCQmedium

A data scientist trains a scikit-learn model with MLflow tracking in a Databricks notebook. They call mlflow.sklearn.log_model(model, 'model') but later find that the Registered Model's schema shows no input signature, preventing automatic schema enforcement during serving. What should they have done to capture the model signature?

A.Register the model in the workspace Model Registry and then edit the signature in the UI
B.Call mlflow.sklearn.log_model(model, 'model', signature=mlflow.models.infer_signature(X_train, model.predict(X_train)))
C.Log the training DataFrame as an artifact and add a signature manually in the MLmodel file after training
D.Enable autologging with mlflow.sklearn.autolog() before training and logging
AnswerB

mlflow.models.infer_signature(X_train, model.predict(X_train)) examines the training DataFrame and the model's predictions to build an MLflow ModelSignature containing input and output column types. Passing it to log_model persists the signature in MLmodel, enabling schema validation and automatic enforcement by Model Serving. It is the documented, lightweight way to capture schema without manual specification.

Why this answer

The model signature is created and attached during logging, not afterward. mlflow.models.infer_signature derives input and output schema from sample data and predictions, and passing it to log_model stores it in the MLmodel file. This enables schema enforcement and automatic serving input validation. Other options either do not create a signature or attempt unsupported post-hoc edits.

Exam trap

The trap here is assuming that autologging or the Model Registry UI can add a signature after the fact, when signatures must be captured at log time.

7
Multi-Selectmedium

Which TWO factors should be considered when choosing between Batch Inference and Real-time Inference for a model in Databricks?

Select 2 answers
A.The memory limit of the driver node during the model training phase.
B.The required latency for the model predictions to reach the end user.
C.The total volume and frequency of the incoming data requests.
D.The number of features used in the machine learning model.
E.The programming language used to write the ML model training code.
AnswersB, C

Latency is the primary driver for choosing between real-time and batch inference. If the user requires sub-second responses to an event, real-time serving is necessary. If the predictions can be processed asynchronously, batch inference is significantly more efficient and cost-effective for handling high volumes of data.

Why this answer

Choosing between batch and real-time inference depends on latency requirements and data availability. Batch is suited for high-volume, non-time-sensitive tasks, whereas real-time is necessary for low-latency, event-driven predictions. Understanding these trade-offs is essential for ML architects to design cost-effective and performant serving strategies that align with business needs, as they dictate the infrastructure configuration, monitoring complexity, and the overall deployment strategy for the model lifecycle.

Exam trap

Candidates often over-focus on model accuracy or training speed, forgetting that deployment strategy is primarily dictated by the business requirements of latency and data throughput volume rather than model complexity.

8
MCQmedium

A team is transitioning from local model development to production in Databricks. Which THREE practices should be implemented to ensure a successful MLOps workflow?

A.Always register models in the MLflow Model Registry.
B.Keep all training and production code in a single notebook for simplicity.
C.Implement automated CI/CD pipelines for deployment.
D.Monitor model performance and trigger retraining when drift is detected.
E.Allow all team members to have 'Admin' access to production models.
AnswerA, C, D

Registering models provides a centralized, version-controlled repository. This allows for clear tracking of which model version is in which stage, enables audit logs of changes, and facilitates integration with deployment tools, which is necessary for managing model provenance and ensuring only approved versions are deployed to production endpoints.

Why this answer

Moving to production requires rigorous automation and governance. Using a centralized model registry ensures version control and auditability. Automated CI/CD pipelines ensure that model training and deployment are consistent and repeatable, reducing the risk of manual errors.

Finally, monitoring model performance in production is critical to detect drift and trigger retraining, ensuring the model remains accurate and relevant as the underlying data distribution changes over time in the real-world environment.

Exam trap

Candidates often focus only on model training, forgetting that production MLOps requires a holistic approach including continuous monitoring for drift and automated CI/CD deployment pipelines.

9
MCQmedium

An ML engineer is developing a model in a Databricks notebook. They want to log the model and its dependencies to MLflow, ensuring that the exact library versions used during training are captured. They also need to register the model in the Databricks Model Registry. Which approach should they use?

A.Save the model using joblib.dump and then use MLflow to log the file as an artifact.
B.Use mlflow.pyfunc.log_model to log the model, then call mlflow.register_model to register it.
C.Log the model using mlflow.log_artifact and then manually register it via the MLflow UI.
D.Log the model using mlflow.sklearn.log_model with the registered_model_name parameter set.
AnswerD

Using mlflow.sklearn.log_model with registered_model_name logs the model and automatically registers it in the Databricks Model Registry. MLflow captures the model's dependencies, including library versions, via the conda environment and requirements.txt, ensuring reproducibility. This approach satisfies both logging and registration in one step, making it the most efficient and correct method.

Why this answer

Logging a model with a framework-specific MLflow flavor function such as mlflow.sklearn.log_model captures the model, its dependencies, and environment details, ensuring reproducibility. Setting registered_model_name registers the model in the Databricks Model Registry in the same operation, streamlining the workflow. Other methods either lack dependency capture or require additional manual steps.

Exam trap

The trap here is assuming that any artifact logging method will automatically capture dependencies and register the model, when only model logging functions with registered_model_name provide that integrated behavior.

10
MCQhard

When using the Databricks Feature Store for real-time inference, why is it recommended to use a specific online store (e.g., Cosmos DB) instead of querying the offline store?

A.The offline store is not compatible with the Python programming language.
B.The offline store cannot handle concurrent requests from multiple inference endpoints.
C.Online stores are designed for low-latency random access lookups.
D.The offline store requires a GPU cluster to perform any queries.
AnswerC

Real-time inference demands rapid, single-record retrieval. Online stores provide this by using indexing and caching mechanisms that allow for millisecond-level latency. Offline stores, by contrast, are designed for batch processing, which makes them unsuitable for servicing individual real-time prediction requests in a production environment.

Why this answer

Offline stores (Delta Lake) are optimized for high-throughput, sequential read/write operations common in batch training, but they suffer from high latency during random access lookups. Online stores are optimized for low-latency point lookups (Key-Value access). Separating these storage tiers allows ML models to achieve the sub-millisecond response times required for real-time applications while still leveraging the massive analytical capabilities of the offline store for training.

Exam trap

Candidates often assume the offline store can be used for real-time inference because it contains all the data, ignoring the high-latency performance limitations of Delta Lake.

11
MCQmedium

An ML engineer is building a recurring batch inference pipeline in Databricks. The model is registered in Unity Catalog as a model version and must always use the version currently tagged as 'champion', which is reassigned after each retraining run. The engineer wants the inference notebook to resolve this alias at runtime rather than hard-coding a version number. Which approach should the engineer use in the notebook?

A.Load the model with mlflow.pyfunc.load_model("models:/catalog.schema.model_name@champion") to resolve the alias at runtime.
B.Load the model with mlflow.pyfunc.load_model("models:/catalog.schema.model_name/Production") to resolve the stage at runtime.
C.Read the model version number from the MLflow Tracking API at runtime and pass it to mlflow.pyfunc.load_model with a numeric version URI.
D.Download the model artifacts from the Unity Catalog volume path with dbutils.fs.cp and load them with mlflow.pyfunc.load_model using the local path.
AnswerA

The models:/ URI with the @champion suffix resolves the Unity Catalog model alias to the currently assigned version at load time, so the notebook always picks up the latest promoted model without edits. This is the documented pattern for alias-based deployment in Databricks, and it works with mlflow.pyfunc.load_model as well as scoring inside the notebook or a job task.

Why this answer

Using a model alias in the models:/ URI lets the pipeline resolve the promoted version at runtime, so retraining and promotion do not require editing the inference notebook. The alias is a mutable pointer managed in Unity Catalog, and mlflow.pyfunc.load_model understands the models:/ scheme with the @alias suffix, making it the intended mechanism for champion-based batch scoring.

Exam trap

The trap here is assuming that MLflow model stages like Production still apply inside Unity Catalog, when Unity Catalog models rely on aliases instead.

12
MCQmedium

An ML engineer is building a Databricks Job that trains a model with scikit-learn and needs to capture hyperparameters, evaluation metrics, and the resulting model artifact for each run. The team wants to compare runs visually in the workspace and later promote the best model to the Model Registry. Which MLflow capability should the engineer use to record these per-run details?

A.Databricks Feature Store with online store enabled
B.MLflow Tracking with autologging enabled for scikit-learn
C.Databricks Jobs task dependencies with a condition task
D.MLflow Model Registry stage transitions
AnswerB

MLflow Tracking records parameters, metrics, tags, and artifacts for each run, and the scikit-learn autologging integration automatically captures estimator hyperparameters, evaluation metrics, and the fitted model artifact. This satisfies the need to compare runs in the workspace UI and gives a model artifact that can be registered afterward, with no manual logging calls required for standard estimators.

Why this answer

Recording per-run parameters, metrics, and model artifacts is the job of MLflow Tracking, and the scikit-learn autologging integration captures these automatically for standard estimators. That gives the team comparable runs in the workspace experiment UI and a logged model artifact that can subsequently be registered and promoted through the Model Registry.

Exam trap

The trap here is assuming that the Model Registry or Feature Store performs experiment tracking, when in fact only MLflow Tracking logs the per-run parameters, metrics, and artifacts.

13
MCQmedium

A data scientist is building a training pipeline where raw event data lands in a Delta table. They need to transform the data, train a model, and register it to the Databricks Model Registry. The pipeline must run daily on a schedule and send an email alert if training fails. Which Databricks construct should they use to orchestrate the entire workflow?

A.A Databricks Job with multiple tasks that include a notebook task for data transformation, a notebook task for training, and a task that registers the model, with email notifications configured on the Job.
B.An MLflow Project that defines the transformation, training, and registration steps in its MLproject file and is run with mlflow run on a schedule.
C.A Delta Live Tables pipeline that performs transformations and automatically trains and registers the model as part of the pipeline.
D.A Databricks SQL dashboard that queries the raw Delta table, calls an external training service via a webhook, and emails the results.
AnswerA

A Databricks Job supports multiple tasks with dependencies, so transformation, training, and registration can be sequenced in one workflow. Job-level email notifications on failure satisfy the alerting requirement, and the Job scheduler runs it daily. This is the native Databricks orchestration construct for multi-step ML workflows.

Why this answer

Databricks Jobs with multiple tasks are the native way to orchestrate multi-step ML workflows. You can define task dependencies so transformation runs before training, and training before registration. Job-level notifications send email alerts on failure, and the built-in scheduler handles the daily cadence.

This satisfies all requirements without external tooling.

Exam trap

The trap here is assuming MLflow Projects provide scheduling and alerting, when they only package code and must be orchestrated by a separate scheduler.

14
MCQmedium

Your team is experiencing 'data drift' in production where the model's accuracy drops over time. What is the most recommended Databricks-native approach to address this?

A.Increase the size of the serving cluster to handle more data points.
B.Implement a retraining pipeline that is triggered when performance metrics drop.
C.Hard-code the input feature ranges in the inference function to filter out outliers.
D.Switch to a more complex model architecture to better fit the production data.
AnswerB

A robust ML workflow includes continuous monitoring and an automated retraining loop. When performance metrics drop, a pipeline should be triggered to retrain the model on the most recent data. This effectively mitigates data drift by keeping the model updated with the current characteristics of the production environment.

Why this answer

Addressing data drift involves monitoring incoming production data and comparing it against the training data distribution. By using Delta Lake's time-travel capabilities and MLflow's experiment tracking, teams can identify when and why model performance degrades. This proactive monitoring allows for timely retraining of the model with updated data, ensuring that production predictions remain accurate and aligned with the current real-world environment, which is vital for long-term model reliability.

Exam trap

Candidates may suggest manual retraining or ignoring the issue, failing to recognize that a systematic, data-driven approach using monitoring and automated triggers is the standard Databricks recommendation.

15
MCQmedium

When designing a production-grade machine learning workflow, which THREE of the following are necessary to ensure the pipeline is observable and recoverable?

A.Use Delta Lake to version the training datasets.
B.Configure the job to automatically ignore all failures and continue.
C.Implement model checkpointing during the training process.
D.Log all training metrics and parameters to MLflow.
E.Run all jobs manually to avoid the risk of automation errors.
AnswerA, C, D

Delta Lake's time travel and versioning capabilities allow you to access the state of your data at any point in time. This is critical for reproducing training results and debugging data-related issues, as it ensures you can consistently train models on the exact same data as the original run.

Why this answer

Observability and recoverability are pillars of reliable MLOps. Comprehensive logging (via MLflow) provides the 'why' behind a model's state. Versioning datasets ensures you can recreate the input data if a model fails.

Finally, robust checkpointing during training allows the job to resume from the last successful state rather than restarting from scratch, saving significant time and resources when failures occur in large, long-running training jobs.

Exam trap

Test-takers frequently select general performance metrics like accuracy tuning as pillars of operational recoverability, confusing model optimization with infrastructure observability and job failure recovery mechanisms.

16
Multi-Selectmedium

An ML engineer is setting up a Databricks Job to retrain a production model nightly. The job must run only when upstream data validation succeeds, notify the team on failure, and avoid retraining when the input data has not changed. Which TWO capabilities of Databricks Jobs directly support these requirements? (Choose two.)

Select 2 answers
A.Cluster autoscaling to add workers when the training dataset grows
B.Job-level and task-level notifications that alert on failure or duration thresholds
C.Job compute configured with photon acceleration for the training tasks
D.Task dependencies with condition tasks to branch based on upstream task outcomes
E.Repair runs to rerun only failed tasks after a partial failure
AnswersB, D

Databricks Jobs support email, webhook, and other notifications configured per task or for the whole job, triggered on failure, success, or duration limits. Configuring failure notifications satisfies the requirement to inform the team when the nightly retraining pipeline breaks, complementing the branching logic that controls whether training runs at all.

Why this answer

Condition tasks provide the branching needed to run training only when validation passes and data has changed, while task and job notifications deliver the failure alerts. Together they implement the conditional retraining and notification requirements, whereas autoscaling, repair runs, and Photon address performance or recovery rather than control flow and alerting.

Exam trap

The trap here is treating performance features such as autoscaling or Photon as workflow controls, when conditional execution and alerting are handled by condition tasks and notifications.

17
MCQeasy

A data scientist wants to compare the performance of three hyperparameter configurations for a Spark ML model. They need a central place to view metrics like RMSE and MAE across runs, and to filter runs by parameters. Which Databricks capability should they use?

A.Databricks Jobs with task dependencies
B.MLflow Tracking with the Databricks-hosted experiment UI
C.Databricks Feature Store
D.Databricks Model Registry
AnswerB

MLflow Tracking records parameters, metrics, and artifacts for each run. The Databricks-hosted MLflow experiment UI provides side-by-side run comparison, metric charts, and filtering by parameters. This directly satisfies the need to compare hyperparameter configurations using RMSE and MAE in a central place.

Why this answer

MLflow Tracking, accessed through the Databricks experiment UI, is the component that logs and visualizes run metrics, parameters, and artifacts. It enables filtering and side-by-side comparison of runs, which is exactly what the data scientist needs for evaluating hyperparameter configurations using RMSE and MAE.

Exam trap

The trap here is confusing the Model Registry or Jobs with the experiment tracking UI, when only MLflow Tracking provides run comparison and metric visualization.

18
MCQeasy

A data scientist finishes a training notebook and wants to capture the source code revision and the Git repository URL on the MLflow run so reviewers can reproduce the exact code state. The repository is already connected to the Databricks workspace through Git integration. Which mechanism records this information automatically?

A.Enabling autologging, which always injects Git metadata regardless of where the notebook is executed.
B.MLflow Git-based source tracking, which stores the repository URL and commit hash as run tags when the notebook runs from a Git folder.
C.Manually calling mlflow.set_tag() with a hand-copied commit hash after each training run.
D.Registering the model in Unity Catalog, which copies the workspace Git configuration into model version metadata.
AnswerB

When a notebook executes from a Databricks Git folder, MLflow records the Git repository URL and commit hash as run tags such as mlflow.source.git.repoURL and mlflow.source.git.commit. Reviewers can then check out that exact revision to reproduce the run. No manual logging call is needed, which is why this is the correct mechanism for the described requirement.

Why this answer

Databricks integrates Git folders with MLflow so that runs executed from a connected repository automatically receive source tags including the repository URL and commit hash. This gives reviewers a precise revision to check out for reproduction. Manual tagging is not automatic, autologging targets framework artifacts rather than Git metadata, and model registration does not populate run source tags.

Exam trap

The trap here is assuming autologging captures Git provenance, when Git source tags come from running the notebook inside a connected Git folder.

19
MCQmedium

A data scientist is tuning a scikit-learn random forest with Hyperopt in a Databricks notebook. Each trial trains on a 40 GB Delta table, and the scientist notices that every trial re-reads the full table from cloud storage, making the search slow. Which change best accelerates the hyperparameter search while preserving correctness?

A.Switch the search algorithm from the default TPE to Random Search to reduce per-trial overhead.
B.Enable adaptive query execution and rely on Spark to skip re-reading unchanged files automatically.
C.Load the Delta table into a Spark DataFrame once, cache it, and have each trial convert the cached data to pandas.
D.Increase max_evals so more trials run and the overhead is amortized.
AnswerC

Caching the Spark DataFrame materializes the data in cluster memory or local disk so subsequent trials reuse it instead of re-reading cloud storage. Each trial still converts the cached data to the pandas representation scikit-learn needs, preserving the exact training data and therefore correctness. This removes the dominant repeated I/O cost and is the standard way to speed up Hyperopt loops on Databricks.

Why this answer

Hyperopt runs many independent trials, and when each trial re-reads a large Delta table, storage I/O dominates runtime. Caching the Spark DataFrame once outside the objective function materializes the data so trials reuse it, while converting to pandas inside each trial keeps scikit-learn training identical. This preserves data and model semantics and removes the repeated read that made the search slow.

Exam trap

The trap here is tuning the search algorithm or trial count when the real bottleneck is repeated data reads that caching would eliminate.

20
Multi-Selecthard

An ML platform team is standardizing how models move from experimentation to production on Databricks. They want promotion decisions to be auditable and to prevent unvalidated models from serving live traffic. Which TWO practices align with Databricks Model Registry and Unity Catalog governance? (Choose two.)

Select 2 answers
A.Embed promotion logic entirely in notebook widgets so reviewers approve versions from the notebook UI.
B.Allow any team member with workspace access to register and promote models to reduce bottlenecks.
C.Rely on MLflow stages like Staging and Production to gate which model version receives live traffic.
D.Store models in Unity Catalog with a three-level name and control access using catalog, schema, and model privileges.
E.Use model aliases such as champion and challenger to mark the versions approved for serving and evaluation.
AnswersD, E

Unity Catalog governs models as securable objects, so privileges such as USE CATALOG, USE SCHEMA, and EXECUTE on the model determine who can load or serve them. This gives the platform team auditable, centralized access control instead of relying on workspace-local registry permissions, and it prevents unauthorized consumers from loading unvalidated versions in a governed environment.

Why this answer

Auditable promotion in Unity Catalog rests on governed model objects and movable aliases. Three-level names plus catalog, schema, and model privileges control who can read, write, and serve models, while champion and challenger aliases record which version is approved for production or evaluation. Together they let consumers reference a stable name that promotion workflows update, with every change visible in the registry and subject to access control.

Exam trap

The trap here is treating MLflow stages as the promotion mechanism for Unity Catalog models, when governed models use aliases and Unity Catalog privileges instead.

21
MCQhard

Why should you use a dedicated Feature Store for your ML workflows instead of just storing features as Delta tables?

A.It makes Delta tables run significantly faster during training.
B.It automatically cleans missing data in the raw tables.
C.It ensures consistency by linking features across training and inference.
D.It is the only way to store data in the Databricks platform.
AnswerC

The Feature Store provides a unified interface that ensures the same features used during training are accessed during serving. This consistency is the most important factor in avoiding training-serving skew, a common and difficult-to-debug issue where models perform perfectly in the lab but fail silently in production due to input differences.

Why this answer

While Delta tables provide high-performance storage, a dedicated Feature Store adds a metadata layer that tracks feature lineage, usage, and schema definitions. It ensures that the exact same feature logic is shared between training and inference pipelines. This consistency is critical for preventing training-serving skew, which is a major, often silent, cause of model performance degradation when moving from development to real-world production environments.

Exam trap

Many candidates believe standard Delta tables automatically handle feature reuse and lineage tracking across training and inference, missing the specific metadata consistency guarantees provided by a Feature Store.

22
MCQmedium

When designing a feature engineering pipeline in Databricks, why should you use the Feature Store instead of standard Delta tables?

A.Delta tables do not support versioning of data.
B.Feature Store provides automatic point-in-time joins to prevent data leakage.
C.Delta tables require external storage, whereas Feature Store is local-only.
D.Feature Store is the only way to perform SQL queries in Databricks.
AnswerB

Point-in-time joins are critical for preventing data leakage during model training. The Feature Store automatically handles these joins, ensuring that features are computed using only data available at the time of the event, which is essential for accurate model evaluation and performance in real-world scenarios.

Why this answer

Feature Store enables feature reuse and point-in-time correctness. Standard Delta tables do not automatically manage feature metadata or linkage between training data and inference features. By using the Feature Store, teams prevent training-serving skew, a common cause of model degradation.

It also provides a centralized catalog for discovery, which improves collaboration and prevents duplicate engineering efforts across different machine learning projects within the organization.

Exam trap

Candidates often assume Feature Store is just for storage, missing the critical point that its primary advantage is point-in-time joins to prevent data leakage during training.

23
MCQmedium

Which Databricks feature allows you to monitor and manage the lineage of data from the source to the final model prediction?

A.MLflow Tracking
B.Unity Catalog
C.Databricks Delta Live Tables
D.Databricks Model Serving
AnswerB

Unity Catalog offers end-to-end lineage tracking, showing the movement and transformation of data across the platform. It maps the dependencies between upstream datasets, feature tables, and downstream machine learning models, providing a complete audit trail that is essential for governance, impact analysis, and troubleshooting production ML pipelines.

Why this answer

Unity Catalog provides comprehensive data lineage capabilities that track how data flows through tables, jobs, and models. By visualizing this lineage, data scientists and engineers can understand the dependencies between raw data, feature tables, and trained models. This is crucial for debugging, impact analysis, and meeting compliance requirements, as it allows users to trace a model's input data back to its origins within the Databricks ecosystem.

Exam trap

Candidates often confuse MLflow tracking with Unity Catalog, assuming that experiment logs provide full data lineage, whereas Unity Catalog is the specific tool designed for end-to-end data flow tracking.

24
MCQeasy

Which of the following is an advantage of using Delta Lake for the data storage layer in an ML workflow?

A.It automatically cleans up old model versions from the MLflow registry.
B.It enables time travel to reproduce historical training data sets.
C.It eliminates the need for feature engineering by auto-generating features.
D.It ensures that the model training code runs on a single node.
AnswerB

Delta Lake's time-travel feature allows users to query data as it existed at a specific point in time. This is critical for reproducing training experiments, as it ensures that the training dataset is identical to the one used in previous runs, regardless of ongoing data updates or transformations.

Why this answer

Delta Lake provides ACID transactions and time-travel capabilities, which are essential for ML workflows. Being able to access previous versions of data (time travel) allows data scientists to reproduce exact training results, ensuring consistency across experiments. This reliability is foundational for auditing and maintaining high-quality models, as it guarantees that the data used for training is well-versioned, consistent, and resilient to failures during concurrent read/write operations common in big data environments.

Exam trap

Candidates often confuse Delta Lake's time-travel feature with model versioning, failing to recognize that time-travel specifically preserves historical training datasets.

25
MCQmedium

A data scientist trains a scikit-learn model in a Databricks notebook and logs it with MLflow. She now needs to promote the exact model artifact to the Databricks Model Registry so that it can be served. Which MLflow API call accomplishes this promotion?

A.mlflow.tracking.MlflowClient().create_registered_model(name)
B.mlflow.register_model(model_uri, name)
C.mlflow.sklearn.save_model(model, path)
D.mlflow.log_artifact(model, artifact_path='model')
AnswerB

mlflow.register_model() takes the run-relative URI of the logged model (for example runs:/<run_id>/model) and a registered model name, then creates or updates the registered model entry in the workspace Model Registry. It is the canonical way to promote an artifact produced by mlflow.sklearn.log_model into the registry, and it returns a ModelVersion object for later stage or alias transitions.

Why this answer

Registering a logged model means creating a named entry in the Model Registry that points at a specific run artifact. The mlflow.register_model() function performs exactly that linkage by accepting the run-relative model URI and the desired registered model name, returning a version that can then be transitioned to a stage or alias for serving. Saving, creating an empty registered model, or logging an artifact each perform only one piece of the workflow and leave no registered version.

Exam trap

The trap here is assuming that logging or saving a model artifact automatically creates a registry entry, when registration is a distinct API call that links an artifact URI to a registered model name.

26
MCQmedium

A data science team is using Databricks Jobs to orchestrate a machine learning pipeline. The pipeline includes a task that trains a model and a subsequent task that evaluates the model. The evaluation task must access the model version produced by the training task. Which mechanism should the team use to pass the model version between tasks?

A.Write the model version to a Delta table and have the evaluation task read it.
B.Store the model version in a file in DBFS and have the evaluation task read it.
C.Use MLflow's tracking API to query the latest model version in the evaluation task.
D.Use task values to pass the model version as a string between tasks.
AnswerD

Databricks Jobs support task values, which allow tasks to pass small amounts of data (like a model version number) to downstream tasks. The training task can set a task value with the model version, and the evaluation task can retrieve it using dbutils.jobs.taskValues.get(). This is the recommended way to pass dynamic values between tasks in a workflow.

Why this answer

Task values are the native Databricks Jobs feature for passing small data between tasks. The training task can set a task value with the model version, and the evaluation task retrieves it using dbutils.jobs.taskValues.get(). This avoids ambiguity and external storage, ensuring the correct version is used.

Exam trap

The trap here is assuming that querying the latest model version is equivalent to the version produced by the training task, which can be incorrect if other runs occur.

27
Multi-Selecthard

Which THREE of the following are benefits of using Databricks Workflows for ML model retraining?

Select 3 answers
A.Integration with MLflow to track parameters and metrics for each automated run.
B.Native support for triggering jobs based on data arrival in specific locations.
C.Built-in automated hyperparameter tuning within the workflow orchestrator itself.
D.Capability to send notifications upon success or failure of the pipeline.
E.Automatic translation of Python code into SQL for faster execution.
AnswersA, B, D

Workflows seamlessly integrates with MLflow, ensuring that every automated retraining job generates a traceable experiment record. This allows teams to analyze the performance of new iterations against historical baselines, which is essential for maintaining model quality and debugging failures in automated production pipelines over long periods.

Why this answer

Databricks Workflows provides a robust orchestration layer to automate the entire ML lifecycle. By automating retraining, teams ensure models remain relevant as data changes over time. Reliability features like retries and failure alerts reduce the operational burden on data scientists, allowing them to focus on model improvement while the infrastructure ensures that production pipelines are consistently refreshed without manual intervention.

Exam trap

Candidates often incorrectly select 'automatic model retraining' as a native feature, whereas Workflows provides the orchestration to trigger retraining, not the automated model logic itself.

28
Multi-Selectmedium

A team is using Databricks Feature Store to create a feature table that will be used for both batch training and online inference. They need to ensure the feature table supports point-in-time lookups for training and low-latency reads for serving. Which TWO of the following statements are correct about meeting these requirements? (Choose two.)

Select 2 answers
A.Point-in-time lookups require the feature table to be partitioned by the timestamp key.
B.The feature table must be published to an online store to support low-latency online inference.
C.Online stores are automatically created when a feature table is registered, requiring no additional configuration.
D.Point-in-time lookups are automatically handled by the feature table's primary key and timestamp key during training.
E.The feature table must be stored in Parquet format to enable low-latency reads.
AnswersB, D

Databricks Feature Store can publish feature tables to an online store (e.g., DynamoDB, Cosmos DB, or online tables) to serve features with low latency. Without publishing, the feature table is only available for batch reads from Delta, which is not suitable for real-time serving. Publishing creates a synchronized online copy that Model Serving can query.

Why this answer

To support both batch training and online inference, the feature table must be published to an online store for low-latency reads, and point-in-time lookups are handled using the primary key and timestamp key during training set creation. Parquet storage and automatic online store creation are not requirements, and partitioning is optional.

Exam trap

The trap here is assuming that online stores are created automatically or that a specific file format is required for low-latency reads, when publishing is an explicit step.

29
MCQmedium

What is the primary advantage of using Databricks Model Serving over deploying a model on a standalone web server?

A.It allows you to write the model logic in any programming language including C++.
B.It provides managed, auto-scaling infrastructure with built-in model versioning.
C.It is cheaper than a dedicated server regardless of the model usage patterns.
D.It removes the need to perform any data preprocessing before inference.
AnswerB

Managed serving provides auto-scaling and high availability without the user needing to manage underlying Kubernetes clusters or server infrastructure. Integration with the Model Registry ensures version consistency, simplifying the deployment pipeline and ensuring that production endpoints are reliable, secure, and always running the latest validated model version.

Why this answer

Databricks Model Serving provides a managed, serverless infrastructure that scales automatically based on load and provides low-latency inference. It integrates directly with the MLflow Model Registry, ensuring that the model version deployed is exactly the one tested. This managed approach handles complex infrastructure concerns like auto-scaling, high availability, and authentication, reducing operational overhead compared to manual deployments on standalone servers, which require custom management of dependencies, security, and scaling.

Exam trap

Candidates often think Model Serving is only about low latency, failing to recognize that the primary benefit is the managed, serverless auto-scaling infrastructure integrated with MLflow.

30
MCQmedium

A data scientist wants to share an MLflow experiment with a teammate. What is the most direct way to ensure the teammate can access the metrics and parameters?

A.Export the experiment data as a CSV and email it to the teammate.
B.Move the experiment notebook to the teammate's private folder.
C.Update the 'Permissions' for the experiment in the Databricks UI.
D.Hardcode the teammate's user ID into the logging script.
AnswerC

Databricks provides a native interface to manage access to experiments. Updating permissions is the correct, secure, and professional way to share experiment results. It allows team members to see all logs, metrics, and parameters within the platform, fostering effective collaboration and ensuring that everyone is working from the same source of truth.

Why this answer

MLflow Experiments are stored within the Databricks workspace and follow standard permission models. By adjusting the 'Permissions' settings on the experiment, the data scientist can grant 'Can View' or 'Can Manage' access to the teammate. This centralized access management ensures that team members can collaboratively review experiment results, compare performance metrics across different runs, and verify the reproducibility of models without needing to manually share raw files or screenshots.

Exam trap

Many students suggest exporting notebooks or downloading CSVs of metrics to share results, ignoring the direct permission management capabilities available natively within the Databricks UI.

31
MCQhard

An ML engineer needs to deploy a model for real-time inference with automatic scaling and a REST endpoint that requires token-based authentication. The model artifacts are already registered in the Databricks Model Registry. Which Databricks capability should be used?

A.MLflow's built-in model serving command, which deploys the model to a local REST server on the cluster and exposes it through the workspace URL.
B.A Databricks Job with a continuous trigger that runs a notebook hosting a Flask application on the driver node.
C.Databricks Model Serving, which creates a scalable REST endpoint for registered model versions with built-in authentication and autoscaling.
D.Databricks SQL endpoint with an AI function that calls the registered model for each row in a query.
AnswerC

Databricks Model Serving is designed to expose registered model versions as REST endpoints, handling autoscaling, availability, and authentication through Databricks tokens or service principals. It integrates directly with the Model Registry, so the engineer can enable serving on a version and get a secured endpoint without managing infrastructure.

Why this answer

Databricks Model Serving provides managed, autoscaling REST endpoints for registered model versions, with authentication handled through Databricks tokens or service principals. It is the native capability for real-time inference and requires no custom Flask hosting, cluster networking, or manual scaling logic, unlike the other options.

Exam trap

The trap here is treating a continuous job or MLflow's local serve command as production serving, when only Model Serving provides managed autoscaling and token-secured endpoints.

32
MCQeasy

Which of the following is a core characteristic of 'Model Serving' in Databricks?

A.It is used strictly for batch processing of large datasets.
B.It automatically scales to handle incoming request volume.
C.It requires the user to manually manage Kubernetes clusters.
D.It stores training data for the model to use at runtime.
AnswerB

Model serving is built to be auto-scaling. It monitors incoming traffic and adjusts the number of instances accordingly, ensuring that the endpoint remains responsive during spikes in demand while optimizing costs during periods of low activity, which is essential for maintaining a reliable production-grade machine learning service.

Why this answer

Model Serving in Databricks allows you to deploy models from the Model Registry as low-latency REST endpoints. These endpoints automatically scale based on demand and handle the underlying infrastructure. This enables data scientists to easily expose their models as services for applications, ensuring high availability and consistent performance without needing to manage the complexities of server configuration, manual load balancing, or capacity provisioning.

Exam trap

Students often think manual cluster scaling or writing custom load-balancing scripts is required for production deployments, forgetting that Databricks Model Serving handles scaling automatically.

33
MCQmedium

An ML engineer is orchestrating an end-to-end machine learning pipeline using Databricks Jobs. The pipeline consists of data preparation, distributed hyperparameter tuning with Hyperopt, and model registration. The engineer needs to pass the best-performing model's run ID from the Hyperopt task to the subsequent model registration task dynamically. Which mechanism should the engineer use in Databricks Jobs to achieve this?

A.Store the run ID in a shared Delta table inside the workspace and query it from the downstream task.
B.Write the run ID to a text file in DBFS root and read it back in the downstream task.
C.Use dbutils.jobs.taskValues to set the run ID in the producer task and get the value in the consumer task.
D.Pass the run ID as a static parameter defined in the JSON job configuration before execution begins.
AnswerC

Task values provide a robust, native API to set key-value pairs in one workflow task and retrieve them in subsequent tasks. This avoids external storage overhead and integrates directly into the Databricks Jobs execution context and UI.

Why this answer

Databricks Jobs task values allow tasks to pass small amounts of data, such as strings or numeric metrics, downstream to other tasks within the same workflow. Using dbutils.jobs.taskValues.set inside the Hyperopt notebook and dbutils.jobs.taskValues.get in the registration notebook enables seamless dynamic orchestration without relying on external storage solutions.

Exam trap

Candidates often assume they need to read and write temporary state files to DBFS or external cloud storage between tasks, forgetting that native task values are specifically engineered for this exact inter-task communication pattern.

34
Multi-Selecthard

An ML engineer is using Databricks Jobs to orchestrate a machine learning pipeline that includes data ingestion, feature engineering, model training, and batch scoring. The engineer wants to ensure that the pipeline is reproducible, handles failures gracefully, and allows for easy debugging of individual tasks. Which TWO features of Databricks Jobs should the engineer leverage to meet these requirements? (Choose two.)

Select 2 answers
A.Notebook parameters to pass dynamic values between tasks.
B.Git integration to version-control notebooks and reference specific commits.
C.Repair runs to re-execute only failed or skipped tasks without rerunning successful ones.
D.Task dependencies to define the order of execution and enable conditional branching.
E.Job clusters that are created for each run and terminated upon completion.
AnswersC, D

Repair runs allow you to rerun only the tasks that failed or were skipped, preserving the results of successful tasks. This is essential for graceful failure handling and efficient debugging, as the engineer can fix an issue and repair the run without redoing the entire pipeline. It directly addresses the requirement to handle failures and debug individual tasks.

Why this answer

Task dependencies define execution order and conditional branching, ensuring that downstream tasks like training run only after upstream tasks succeed and enabling failure handling. Repair runs allow re-executing only failed or skipped tasks, which is critical for debugging and recovering from failures without rerunning the entire pipeline. Together, these Databricks Jobs features provide orchestration, graceful failure handling, and efficient debugging.

Exam trap

The trap here is focusing on code versioning or cluster configuration as the primary means of orchestration and failure recovery, when the core Job features for those needs are dependencies and repair runs.

35
MCQhard

When configuring a Databricks Workflow to automate a machine learning pipeline, which TWO actions are necessary to ensure the pipeline is robust and manageable?

A.Embed all data processing and training logic into a single, massive notebook cell.
B.Use Databricks Workflow tasks to modularize data preparation, training, and evaluation.
C.Configure email notifications to alert team members on job failure or success.
D.Hardcode all file paths and cluster configuration settings directly in the code.
E.Run all pipeline stages on the same shared interactive cluster to save costs.
AnswerB, C

Separating pipeline stages into individual tasks allows for independent execution and retries. This modularity enables developers to monitor each component's success or failure independently, resulting in a cleaner architecture where issues in data cleaning do not require the entire training job to be restarted, improving workflow efficiency significantly.

Why this answer

Robust ML workflows require modularity and error handling. Defining distinct tasks for data preparation, model training, and evaluation allows for granular retries and easier debugging. Furthermore, enabling notifications ensures that failures are addressed promptly.

By using job parameters and dependencies, you create a declarative pipeline that is consistent across environments, reducing the risk of manual configuration errors during deployment cycles and improving overall reliability of the production machine learning system.

Exam trap

Candidates frequently assume that simply running a notebook is sufficient, neglecting the need for modular task separation and automated alerting, which are essential for long-term production reliability and maintenance.

36
MCQeasy

A data scientist wants to compare the accuracy, F1 score, and training duration of several model training runs side by side in a single table, and visually inspect how a hyperparameter affected the metric across runs. Which MLflow capability should be used?

A.The Databricks Jobs UI, where each task run records the metrics emitted by the notebook and displays them in a comparison chart.
B.The MLflow Tracking Server's REST API, which must be queried programmatically to retrieve and compare run metrics.
C.The Databricks Model Registry, where each model version lists its metrics and can be compared against other versions in a table.
D.The MLflow Experiment page, where runs can be selected and compared using the metrics table and chart views.
AnswerD

The MLflow Experiment UI lists all runs with their parameters, metrics, and tags, and supports selecting multiple runs for side-by-side comparison. It also provides chart views that plot a metric against a parameter, which directly answers the need to compare accuracy, F1, and duration and to see hyperparameter effects.

Why this answer

The MLflow Experiment page aggregates all runs with their parameters and metrics, supports selecting multiple runs for a side-by-side table, and offers chart views that plot metrics against parameters. This is the native way to compare accuracy, F1, and training duration and to visualize hyperparameter effects without writing custom code.

Exam trap

The trap here is confusing the Model Registry, which manages versions and stages, with the experiment UI, which is built for comparing runs and metrics.

37
MCQmedium

Which workflow step is essential before promoting a model from 'Staging' to 'Production' in the MLflow Model Registry?

A.Deleting the original training dataset to save storage space.
B.Validating the model performance on a hold-out test set.
C.Manually updating the production database credentials in the script.
D.Restarting the cluster to clear the memory of the training environment.
AnswerB

Validation on a representative hold-out dataset confirms that the model generalizes well to unseen data. This step prevents the deployment of overfitted models that might perform well on training data but fail in production, acting as a critical quality gate that maintains the reliability of the overall MLOps system.

Why this answer

Validation against a held-out test set is a mandatory step before any model promotion. This ensures that the model meets performance requirements and hasn't regressed compared to the currently deployed version. This gated approach, combined with the Model Registry's staging transitions, provides a controlled environment for testing, ensuring that only verified, high-quality models enter the production ecosystem, thus mitigating business risk and maintaining model predictive accuracy in critical applications.

Exam trap

Students often select administrative tasks like registering the model name or creating clusters as the essential step, forgetting that rigorous validation on a hold-out test set is mandatory first.

38
MCQmedium

Refer to the exhibit. Why is the 'signature' parameter included in the log_model call?

A.To encrypt the model artifacts before saving them to DBFS.
B.To enforce input data schema validation during model inference.
C.To specify the hyperparameters used to train the model.
D.To compress the model into a smaller binary file for faster deployment.
AnswerB

Signatures provide a contract between the model and the input data. By logging this schema, the model can automatically check if incoming request payloads match the expected feature names and types. This prevents type-mismatch errors that typically crash inference services when bad data is sent.

Why this answer

The model signature defines the expected schema for inputs and outputs. Providing a signature allows Databricks to perform runtime validation of data types, ensuring that the model receives data in the format it expects. This prevents runtime errors during inference and allows for automated documentation and improved user experience within the MLflow Model Registry, which is crucial for scalable model deployment in production.

Exam trap

Candidates often confuse 'signature' with model versioning or metadata logging, failing to realize its primary function is runtime schema enforcement to prevent type mismatches during inference requests.

39
MCQhard

A team registers a model in Unity Catalog as main.ml.churn_model and wants production scoring jobs to always load the newest approved version without editing job code when a new version is promoted. The team uses the MLflow Python client inside a Databricks job. Which model URI should the scoring code use?

A.runs:/<run_id>/model
B.models:/main.ml.churn_model/Production
C.models:/main.ml.churn_model@champion
D.models:/main.ml.churn_model/3
AnswerC

Unity Catalog model aliases are referenced with the @alias syntax, and the alias is a mutable pointer that promotion workflows update. When the team assigns the champion alias to a new version, the scoring job automatically resolves to that version on its next run without any code change. This satisfies the promotion-without-edit requirement and is the recommended pattern for Unity Catalog models.

Why this answer

Unity Catalog registered models support aliases such as champion, which act as movable pointers to a version. Referencing the model with models:/catalog.schema.model@alias lets consumers load whatever version currently holds the alias, so promotion is a metadata operation rather than a code change. Numeric versions and run IDs are immutable, and stage-based URIs belong to the legacy workspace registry rather than Unity Catalog.

Exam trap

The trap here is carrying over legacy stage names like Production into Unity Catalog URIs, when Unity Catalog models are promoted through aliases instead.

40
MCQmedium

A machine learning engineer is building a training pipeline in Databricks. They want each run to record the exact Git commit hash, the versions of scikit-learn and MLflow used, and the input data path so the run can be reproduced later. Which MLflow tracking capability should they use to capture this information with the least custom code?

A.mlflow.register_model() to create a model version that stores the Git commit hash and data path in its description.
B.mlflow.log_artifact() to upload a text file containing the Git commit hash, library versions, and data path.
C.mlflow.set_tags() combined with MLflow's automatic logging of the source Git commit and environment.
D.mlflow.log_params() to log the Git commit hash, library versions, and data path as run parameters.
AnswerC

MLflow automatically captures the source Git commit hash and the active conda environment when a run starts, and custom metadata such as a data path can be attached with set_tags. Tags are searchable and displayed separately from parameters and metrics, so the engineer gets reproducibility data with almost no code while keeping the run UI clean.

Why this answer

MLflow Tracking automatically records source information, including the Git commit when the run executes from a repository, and the environment details. Tags provide a searchable place for custom metadata such as the input data path. Together they deliver reproducibility context without custom parsing code, unlike parameters, artifacts, or registry entries that are meant for different purposes.

Exam trap

The trap here is assuming that any run metadata must be logged manually as a parameter or artifact, rather than using MLflow's built-in source and environment tracking plus tags.

41
MCQeasy

A data scientist wants to package a training script with its Python dependencies and parameters so that the same code can be rerun on a different Databricks cluster or shared with a colleague and reproduced exactly. Which MLflow component is designed for packaging and reproducing project code in this way?

A.MLflow Projects
B.MLflow Tracking
C.MLflow Models
D.MLflow Model Registry
AnswerA

MLflow Projects package code with an MLproject file that declares entry points, parameters, and environment dependencies such as a conda environment or Docker image. Running the project recreates the specified environment and executes the entry point, so the same code and dependencies can be reproduced on another cluster or shared with a colleague and produce consistent results.

Why this answer

MLflow Projects are the packaging mechanism: an MLproject file names entry points, parameters, and the environment, and running the project recreates that environment to execute the code. This makes training scripts portable and reproducible across clusters and users, which is exactly the requirement described.

Exam trap

The trap here is confusing MLflow Models, which package trained model artifacts, with MLflow Projects, which package the code and environment needed to run training.

42
MCQmedium

A data scientist is training a machine learning model on Databricks and needs to log parameters, metrics, and model artifacts. Which tracking component should be used to ensure the reproducibility of the experiment runs?

A.Databricks Feature Store
B.MLflow Model Registry
C.MLflow Tracking
D.Delta Lake Versioning
AnswerC

MLflow Tracking provides a robust API and UI for logging parameters, code versions, metrics, and output artifacts. It serves as the primary tool for experiment management, allowing users to organize runs into experiments, compare results visually, and maintain a historical audit trail of model development workflows.

Why this answer

MLflow Tracking is the core component for logging parameters, code versions, metrics, and output files when running machine learning code. By recording these artifacts, teams can compare multiple runs, track model lineage, and ensure that experiments are reproducible across different environments. Integrating MLflow into the workflow is essential for transitioning from local notebook experimentation to production-ready MLOps pipelines within the Databricks ecosystem.

Exam trap

Candidates select general storage solutions like DBFS or standard cloud buckets, overlooking the dedicated MLflow tracking component designed specifically for reproducibility.

43
MCQmedium

Which technique should be used to prevent data leakage in ML workflows when performing cross-validation on time-series data?

A.Random sampling with replacement across the entire dataset.
B.Standard k-fold cross-validation with 5 folds.
C.Expanding window cross-validation (time-series split).
D.Removing all features that have high correlation with the target variable.
AnswerC

Expanding window validation maintains the temporal order of the data. By training on a chronologically growing window and testing on the subsequent period, you simulate real-world usage where the model predicts the future based on past data, correctly preventing leakage and ensuring reliable evaluation of the model's accuracy.

Why this answer

In time-series forecasting, standard k-fold cross-validation is inappropriate because it would allow the model to 'see' the future during training, leading to overly optimistic performance estimates. Instead, 'time-series split' or 'walk-forward' validation should be used, where the training set only consists of data points that occur chronologically before the validation set. This preserves the temporal order, providing a realistic assessment of the model's performance on unseen future data.

Exam trap

Candidates often apply standard k-fold cross-validation to time-series data, failing to realize that random splitting causes look-ahead bias, which invalidates the model's performance metrics.

44
MCQeasy

A data scientist wants to track the progress of a training script that runs for several hours on a Databricks cluster. The script uses MLflow and needs to record metrics such as loss and accuracy at the end of each epoch so they can be visualized in real time. Which MLflow API call should be used inside the training loop?

A.mlflow.log_param
B.mlflow.set_tag
C.mlflow.log_metric
D.mlflow.log_artifact
AnswerC

mlflow.log_metric logs a numeric value for a named metric and can be called repeatedly with the same metric name, optionally specifying a step. MLflow stores each value with its timestamp and step, enabling the UI to render a live-updating chart of loss and accuracy per epoch. This is the standard way to track training progress in real time.

Why this answer

To record metrics like loss and accuracy at each epoch for real-time visualization, the training loop should call mlflow.log_metric. This API accepts a metric name, a numeric value, and an optional step, and MLflow stores each call as a data point. The MLflow UI plots these points as a time-series chart, allowing progress monitoring during long training runs.

Exam trap

The trap here is confusing parameters, artifacts, and tags with metrics; only mlflow.log_metric produces the numeric time-series data that MLflow charts over steps or time.

45
MCQmedium

You are building a pipeline where a feature table must be updated daily. Which Databricks construct is the most appropriate for orchestrating this periodic feature engineering job?

A.MLflow Experiments
B.Databricks Jobs
C.Unity Catalog
D.MLflow Model Registry
AnswerB

Databricks Jobs provide the scheduling and orchestration logic required to run feature engineering pipelines automatically. By defining a workflow with dependencies and timing, you can guarantee that the Feature Store is updated correctly every day, ensuring downstream models have access to the latest data.

Why this answer

Databricks Workflows (Jobs) are the recommended tool for orchestrating multi-step data processing tasks, including feature engineering pipelines. By scheduling a Job, you can automate the ingestion, transformation, and write operations to the Feature Store. This ensures that features are refreshed consistently, minimizing data drift and enabling reliable model training or batch inference processes that depend on the most recent feature values.

Exam trap

Candidates rely on external cron jobs or standard notebook scheduling instead of leveraging the native orchestration capabilities built into the platform.

46
MCQeasy

A data scientist has trained a model and wants to deploy it for real-time inference with automatic scaling and without managing infrastructure. Which Databricks feature should they use?

A.Databricks SQL endpoint
B.Databricks Model Serving
C.A standalone Flask app on a Databricks cluster
D.MLflow Model Registry
AnswerB

Databricks Model Serving provides a fully managed, serverless endpoint for real-time inference. It automatically scales based on traffic and requires no infrastructure management. You can enable it directly from the Model Registry or via the serving UI, making it the correct choice for this scenario.

Why this answer

Databricks Model Serving is the managed solution for deploying MLflow models as scalable REST endpoints. It handles provisioning, scaling, and monitoring, so the data scientist can focus on the model. It integrates with the Model Registry and supports real-time inference without infrastructure overhead.

Exam trap

The trap here is confusing the Model Registry, which stores model versions, with Model Serving, which actually hosts the model for predictions.

47
MCQeasy

Which Databricks component is specifically designed to manage the full lifecycle of machine learning models, including registration, versioning, and stage transitions?

A.Delta Lake
B.MLflow Model Registry
C.Unity Catalog
D.Databricks Feature Store
AnswerB

The MLflow Model Registry is the designated tool for versioning, deploying, and managing the lifecycle of machine learning models in Databricks. It allows teams to track model lineage, manage stage transitions, and ensure that deployments are consistent, audited, and easily reversible, which is a fundamental requirement for production-grade MLOps pipelines.

Why this answer

The MLflow Model Registry provides a centralized model store, APIs, and UI to collaboratively manage the full lifecycle of MLflow models. It handles versioning, stage transitions (e.g., Staging to Production), and model annotations. This component is essential for operationalizing machine learning, as it provides a single source of truth for model artifacts and ensures that only validated models are promoted to production environments after passing necessary checks.

Exam trap

Candidates often confuse MLflow Tracking with the Model Registry, failing to distinguish between tracking experiments and managing the formal lifecycle of production-ready model versions.

48
MCQeasy

When designing an ML workflow, what is the primary benefit of using MLflow Projects over executing raw scripts?

A.They automatically scale the underlying compute cluster size based on the task.
B.They allow for the automatic versioning and tracking of data snapshots.
C.They facilitate consistent execution by encapsulating environment and dependency definitions.
D.They replace the need for unit testing individual functions within the pipeline.
AnswerC

MLflow Projects use a project specification file to define dependencies, allowing the environment to be recreated reliably. This ensures that the same code runs identically on any environment, which is vital for professional ML workflows where consistency between development, staging, and production environments is mandatory for reliable model results.

Why this answer

MLflow Projects provide a standard format for packaging reusable data science code. By using a YAML-based environment definition, they ensure that the code runs in the exact same environment across different machines. This is critical for ML workflows to guarantee reproducibility, as it eliminates the 'it works on my machine' problem by explicitly defining the dependencies and environment setup required for the project execution in a portable, platform-agnostic manner.

Exam trap

Candidates often mistake MLflow Projects for simple code versioning tools like Git, failing to recognize that their primary purpose is environment and dependency encapsulation to solve reproducibility issues.

49
MCQmedium

You are training a scikit-learn model with MLflow in a Databricks notebook. The model's preprocessing includes a custom Python function that you wrote in the notebook. You need to register the model to the Databricks Model Registry and later deploy it with Model Serving, ensuring the preprocessing is applied automatically at inference. Which approach should you use?

A.Register the model with the sklearn flavor, and configure a pre-processing hook in the Model Serving endpoint configuration.
B.Log the preprocessing function as a separate artifact using mlflow.log_artifact, and reference it from the model's conda environment.
C.Log the model with mlflow.sklearn.log_model, passing the custom preprocessing function in the signature parameter.
D.Wrap the preprocessing and model in an mlflow.pyfunc.PythonModel subclass, and log it with mlflow.pyfunc.log_model.
AnswerD

A custom PythonModel encapsulates the preprocessing code and the trained model in a single artifact. When logged with mlflow.pyfunc.log_model, MLflow serializes the class and its dependencies into the model directory, so Model Serving loads and executes the preprocessing automatically at inference. This is the supported pattern for custom inference logic in Databricks.

Why this answer

To ensure custom preprocessing runs at inference time, the logic must be packaged inside the logged model artifact. Subclassing mlflow.pyfunc.PythonModel and logging with mlflow.pyfunc.log_model bundles the preprocessing code with the trained model, so Model Serving executes it automatically. Other approaches either misuse parameters, store code without wiring it into scoring, or assume unsupported endpoint hooks.

Exam trap

The trap here is assuming that logging a function as an artifact or passing it to the signature parameter makes MLflow execute it during inference, when only code packaged inside the model artifact is actually run.

50
MCQhard

A machine learning team is using Databricks Feature Store to serve features for a real-time model. They have a feature table that is updated daily with new data. To ensure the online store always has the latest feature values for low-latency inference, which approach should they take?

A.Publish the feature table to an online store using the Databricks Feature Store UI or API, and schedule daily updates.
B.Use the Feature Store's automatic online store synchronization by enabling a Delta Live Tables pipeline.
C.Configure the model to query the offline feature table directly during inference, caching results in memory.
D.Create a materialized view of the feature table in Databricks SQL and use that for online serving.
AnswerA

Publishing to an online store (e.g., DynamoDB, Cosmos DB) enables low-latency lookups. Scheduling daily updates ensures the online store is refreshed with the latest feature values from the offline table. This is the standard pattern for real-time serving with Feature Store, as it synchronizes the online store with the offline source.

Why this answer

To serve features in real time with low latency, the feature table must be published to an online store, which is a key-value store optimized for fast lookups. Scheduling daily updates ensures the online store reflects the latest data. Other approaches like querying the offline table or using materialized views introduce latency and are not designed for online serving.

Thus, publishing and scheduling updates is the correct strategy.

Exam trap

The trap here is assuming that Delta Live Tables or materialized views automatically provide online serving; in reality, you must explicitly publish to an online store.

51
MCQhard

A team trains a model with Databricks Feature Store features and logs the training set using feature_lookups. At inference time, they want the model to automatically retrieve the same feature values from the online store so the serving endpoint does not require the caller to supply those features. What must the team do when logging the model so this automatic lookup works?

A.Wrap the model in a PythonModel that calls the Feature Store client inside its predict method
B.Publish the feature table with a primary key and enable streaming on the source Delta table
C.Log the model with the feature engineering client's log_model and include the feature_lookups used during training
D.Register the model in Unity Catalog and grant the serving endpoint EXECUTE on the registered model
AnswerC

The feature engineering client's log_model records the feature lookups alongside the model artifact, so the model metadata knows which feature tables and lookup keys to use. When the model is served with online store access, the endpoint resolves those features itself, meaning callers only provide the primary keys and the model fetches the remaining feature values automatically.

Why this answer

When a model is logged through the feature engineering client with the same feature lookups used during training, the model metadata carries the information needed to resolve those features from the online store at serving time. Callers then only supply primary keys, and the endpoint assembles the full feature vector automatically, keeping training and inference feature logic consistent.

Exam trap

The trap here is believing that registry access grants or feature table publication enable automatic feature retrieval, when only recorded feature lookups in the logged model provide that behavior.

52
MCQeasy

A data scientist is using MLflow Tracking in Databricks to compare multiple runs of a hyperparameter tuning experiment. They want to quickly identify the run with the lowest validation loss and then register that model version in the MLflow Model Registry. Which MLflow UI feature allows sorting runs by a specific metric to find the best run?

A.The runs table, where you can click on the metric column header to sort runs ascending or descending.
B.The experiment notes field, where you can manually record the best run ID.
C.The model registry page, which automatically sorts models by their latest version's metrics.
D.The run comparison view, which displays all runs side-by-side with their metrics.
AnswerA

The runs table in the MLflow UI lists all runs and allows sorting by any metric column. Clicking the column header sorts runs by that metric, enabling quick identification of the run with the lowest validation loss. This is the most efficient way to find the best run for registration.

Why this answer

The runs table in the MLflow UI provides a sortable list of all runs, with columns for parameters and metrics. Clicking a metric column header sorts runs by that metric, making it easy to spot the run with the lowest validation loss. This is the standard workflow for selecting the best run to register.

Exam trap

The trap here is confusing the run comparison view with the runs table, as both show metrics but only the runs table supports sorting.

53
MCQhard

An ML engineer is configuring a Databricks Job to retrain a model daily. The job must run a notebook that reads from a feature table, trains a model, and registers it to the Model Registry. The engineer wants to ensure that the job fails immediately if the model's accuracy drops below a threshold. Which approach should they use?

A.Configure the job to send an email alert if accuracy is below threshold, but allow the job to succeed.
B.Add a task that runs a Python script to query the MLflow API and raise an exception if accuracy is below threshold.
C.Use Databricks Model Serving to monitor the model's accuracy and automatically roll back if it drops.
D.Set a timeout on the training task so that it fails if accuracy is not reached within a time limit.
AnswerB

Adding a task that checks the model's accuracy via MLflow API and raises an exception will cause the job to fail if the threshold is not met. This is a straightforward way to enforce a quality gate within the job's task graph, ensuring that subsequent tasks or the job itself fail.

Why this answer

To make the job fail when accuracy is below a threshold, the engineer should add a task that evaluates the metric and raises an exception. This task can run after training and before registration, ensuring the job stops and no subpar model is registered. This is a common pattern for implementing quality gates in ML pipelines.

Exam trap

The trap here is confusing monitoring or alerting with job failure; only a task that raises an exception will cause the job to fail.

54
Multi-Selectmedium

An ML engineer is configuring a Databricks Job to automate nightly retraining of a model. The job must (1) run only after the upstream feature engineering job succeeds, and (2) notify the team via email if the training task fails. Which TWO configurations satisfy these requirements? (Choose two.)

Select 2 answers
A.Use a schedule with a cron expression that triggers the training task five minutes after the feature engineering job is expected to finish, and enable email alerts on failure.
B.Configure the training task with a depends_on entry referencing the feature engineering task, and add an email notification for the job's failure event.
C.Add a task dependency so the training task depends on the upstream feature engineering task, and configure an email notification on the job for failed runs.
D.Chain the notebooks by calling dbutils.notebook.run from the training notebook to invoke the feature engineering notebook, and rely on the notebook's exception handling to email the team.
E.Set up an MLflow webhook that triggers the training notebook whenever a new model version is registered, and configure the webhook to send email on failure.
AnswersB, C

The depends_on field in a Databricks Job task definition enforces that the referenced upstream task completes successfully before the dependent task runs. Pairing that with a job-level email notification on failure satisfies both the ordering and alerting requirements, making this a correct configuration for the scenario.

Why this answer

Databricks Jobs provide first-class task dependencies via depends_on, which guarantees a downstream task runs only after its upstream task succeeds, and job-level notifications that can email the team on failure. These native features satisfy both the ordering and alerting requirements without custom code or fragile timing assumptions.

Exam trap

The trap here is relying on time-based scheduling or notebook chaining to enforce order, when only a declared task dependency guarantees the upstream job succeeded first.

55
MCQhard

An ML engineer runs an automated hyperparameter sweep with MLflow on a Databricks cluster. The sweep launches 200 runs, and the engineer wants to retrieve, in a notebook, the run ID of the single run that achieved the highest validation accuracy so it can be registered. Which approach correctly identifies that run?

A.Call mlflow.get_run(run_id) for each run created by the sweep and compare the metrics in Python.
B.Use mlflow.list_experiments() and select the experiment with the most runs, then register its latest version.
C.Use mlflow.search_runs() with an experiment ID, order by metrics.validation_accuracy descending, and take the first row's run_id.
D.Read the MLflow experiment's artifact location from DBFS and parse the meta.yaml files to find the highest metric.
AnswerC

search_runs() returns a pandas DataFrame of run metadata and metrics for an experiment. Ordering by the validation accuracy metric in descending order places the best run first, and its run_id column gives the identifier needed for registration. This is the supported programmatic way to query runs and works with the order_by argument using the metrics.<name> syntax.

Why this answer

To find the best run across many, the tracking API must be queried with a metric sort. mlflow.search_runs() supports order_by on metrics.<name> and returns a DataFrame whose first row, when sorted descending, is the top-scoring run. Its run_id can then be passed to registration. Fetching individual runs or reading raw artifact files cannot rank the sweep and relies on unsupported or incomplete data.

Exam trap

The trap here is reaching for get_run() or filesystem parsing to find the best run, when the tracking API's search_runs() is the supported way to sort runs by metric.

56
Multi-Selecthard

Which TWO of the following are benefits of using the Databricks Feature Store for machine learning workflows?

Select 2 answers
A.It automatically scales the compute resources for training deep learning models.
B.It enables consistent feature definitions across training and inference pipelines.
C.It provides point-in-time joins to prevent target leakage.
D.It automatically retrains models when feature data changes.
E.It converts unstructured image data into structured feature vectors.
AnswersB, C

A primary benefit is the ability to share feature definitions between training and online or batch inference. This consistency ensures that the exact same transformations applied to features during training are reproduced during inference, effectively mitigating the common risk of training-serving skew in production systems.

Why this answer

The Databricks Feature Store facilitates feature reuse across different teams and projects, preventing redundant engineering work. Additionally, it ensures point-in-time correctness by preventing data leakage during training through time-travel capabilities. These features are critical for maintaining consistency between training and inference environments, reducing the likelihood of training-serving skew, and streamlining the overall MLOps lifecycle within the Databricks platform.

Exam trap

Candidates mistakenly believe feature stores only speed up training queries, ignoring their critical role in preventing data leakage and ensuring training-serving consistency.

57
MCQhard

A machine learning team is using Databricks Jobs to orchestrate a multi-step ML pipeline: data ingestion, feature engineering, model training, and batch inference. They need to ensure that if the model training step fails, the batch inference step does not run, and that the entire pipeline can be retried from the failed step. Which Databricks Jobs feature should they use to achieve this?

A.Use a single notebook with all steps and rely on the notebook's built-in error handling to skip inference on failure.
B.Set up a continuous job that always runs all tasks, and use conditional logic in the inference task to check if training succeeded.
C.Configure each task with a dependency on the previous task and set the job to repair runs on failure.
D.Use Databricks Jobs with a cluster that has autoscaling enabled and set the inference task to run only if the training task logs a success metric.
AnswerC

Databricks Jobs supports task dependencies, ensuring that a task runs only after its dependencies succeed. The repair run feature allows retrying a failed run from the point of failure, skipping already successful tasks. This meets the requirement of conditional execution and efficient retry.

Why this answer

Databricks Jobs allows defining tasks with dependencies, so the inference task will only run if the training task succeeds. The repair run feature enables retrying a failed run from the failed task, reusing successful task results. This combination provides both conditional execution and efficient recovery.

Exam trap

The trap here is assuming that a single notebook or continuous job can provide the same level of dependency management and repair capabilities as Databricks Jobs with task dependencies.

58
MCQhard

An ML engineer registers a model in Unity Catalog and needs a stable pointer that always resolves to the version currently approved for production, without modifying calling code each time a new version is promoted. Which Model Registry feature should the engineer configure to provide this stable reference?

A.A model version tag recording the approval date
B.A registered model description summarizing the production candidate
C.A model version alias such as 'champion' assigned to the approved version
D.A workspace-level permission grant on the registered model
AnswerC

Aliases are named references that can be reassigned to any model version within a registered model. Serving code can target the alias instead of a version number, so promoting a newly approved version is a matter of moving the alias, and consumers automatically resolve to the newly approved version without any code change or redeployment of callers.

Why this answer

Model version aliases provide exactly the stable, reassignable pointer described: consumers load by alias, and promotion is performed by moving the alias to the newly approved version. Tags, descriptions, and permission grants are metadata or access controls and cannot redirect consumers to a chosen version at load time.

Exam trap

The trap here is assuming that tags or descriptions can be used programmatically to select a version, when only aliases (or stages in older registries) act as resolvable pointers.

59
MCQeasy

What is the primary benefit of using a 'Job Cluster' instead of an 'All-Purpose Cluster' for automated ML workflows?

A.All-purpose clusters are always faster for training large models.
B.Job clusters provide lower costs and better environment isolation.
C.Job clusters allow for interactive debugging of code during runtime.
D.All-purpose clusters are required for scheduling workflows in Databricks.
AnswerB

Job clusters are specifically optimized for automated workflows. By being ephemeral, they reduce costs significantly compared to keeping an all-purpose cluster running. Additionally, because they are isolated from interactive notebooks, they guarantee that the job runs in a predictable environment without side effects from other users or shared session configurations.

Why this answer

Job clusters are ephemeral compute resources that are created for a specific job and terminated immediately after completion. They are significantly more cost-effective because they use specialized pricing and ensure that resources are not idling between runs. Using job clusters also ensures complete environment isolation, preventing interference from other users or interactive processes, which is foundational for ensuring reproducible and reliable machine learning pipeline execution in production environments.

Exam trap

Candidates incorrectly assume all-purpose clusters are better for automated workflows because they remain active, missing that job clusters provide cheaper, ephemeral, and isolated execution.

60
MCQeasy

A data scientist needs to run the same feature-engineering notebook against three different parameter sets in a Databricks Job. She wants each parameter set to execute independently and in parallel, with separate logs. Which Job feature should she use?

A.A job cluster with autoscaling enabled and a single task.
B.A single notebook task with the three parameter sets passed as a JSON string argument.
C.Three separate jobs created from the same notebook, each with its own parameter.
D.A for-each task that iterates over a parameter array and runs the notebook once per element.
AnswerD

A for-each task takes an input array and runs its nested task once per element, optionally in parallel. Each iteration is a separate task run with its own logs and status, which matches the requirement for independent parallel executions with distinct parameter values. The nested notebook receives the current element as a parameter.

Why this answer

The for-each task is designed for iterating a nested task over an input array, producing one task run per element with its own logs and status. That gives the independent, parallel executions the scientist needs while keeping everything in one workflow. Passing a JSON string, creating separate jobs, or relying on autoscaling all fail to produce per-parameter task runs with isolated logs.

Exam trap

The trap here is treating cluster autoscaling or a single parameterized task as a way to run multiple parameter sets, when iteration requires a for-each task over an array.

61
MCQmedium

A data scientist is preparing a feature table in Databricks Feature Store. To ensure the feature table can be used for online inference with low latency, which step is mandatory?

A.Register the feature table in the Unity Catalog without any additional configuration.
B.Call the write_batch method to export features to an S3 bucket for the model to read at runtime.
C.Use the publish_table method to sync features to an online store configured in the Feature Store.
D.Include the feature table in a Databricks Job that runs every minute to keep the data fresh.
AnswerC

The publish_table method is the dedicated API function in Databricks Feature Store designed to replicate feature data from the offline store to an online store. This ensures the features are available for fast key-value lookups, fulfilling the latency requirements necessary for real-time model scoring requests.

Why this answer

To enable online serving, the feature table must be published to a supported online store like Amazon DynamoDB, Azure Cosmos DB, or Google Cloud Bigtable. This process decouples the feature retrieval from the complex logic of feature engineering pipelines, allowing real-time models to fetch pre-computed features in milliseconds rather than recalculating them during the inference request, which is critical for high-throughput production ML applications.

Exam trap

Students often assume registering a model in the Workspace automatically provisions online feature retrieval, missing the explicit sync requirement to an online store.

62
MCQmedium

A data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts and metrics are logged automatically without adding manual logging code to the training script. Which approach should they use?

A.Call mlflow.set_tracking_uri() inside the training function to redirect logs to the workspace.
B.Wrap the training logic within an mlflow.start_run() block without any additional configuration.
C.Execute mlflow.autolog() at the start of the notebook cell before running the model training code.
D.Configure the Databricks cluster environment variable MLFLOW_TRACKING_ENABLED to true.
AnswerC

Calling mlflow.autolog() enables the library-specific hooks that automatically record metrics, parameters, and artifacts during the execution of supported model training methods. This is the standard practice for Databricks ML workflows to reduce manual instrumentation overhead while ensuring that every model iteration is fully documented and tracked automatically.

Why this answer

MLflow provides autologging capabilities that automatically capture parameters, metrics, and models when using popular machine learning libraries like Scikit-learn, TensorFlow, or PyTorch. By calling mlflow.autolog() before the training code, the framework instruments the library calls to record telemetry automatically. This approach minimizes boilerplate code and ensures consistency across experiments, which is essential for auditability and model reproducibility in collaborative Databricks environments, as it prevents human error in manual logging.

Exam trap

Students mistakenly think autologging happens automatically without code invocation, forgetting they must explicitly call 'mlflow.autolog()' within the notebook.

63
MCQeasy

A data scientist is comparing multiple hyperparameter configurations for a model and wants to view the resulting metrics side by side in a single interface, sort runs by accuracy, and drill into individual run details. Which MLflow component provides this capability?

A.MLflow Tracking UI, which lists runs within an experiment and supports sorting and filtering by metrics.
B.MLflow Model Registry UI, which shows model versions and their stage transitions.
C.MLflow Projects dashboard, which lists available projects and their entry points.
D.Databricks Jobs run history, which shows task durations and statuses for scheduled workflows.
AnswerA

The MLflow Tracking UI is the built-in interface for viewing runs in an experiment. It displays metrics, parameters, and tags in a sortable table, allows filtering, and lets users click into a run to see artifacts and details. This directly supports comparing hyperparameter configurations side by side and sorting by accuracy.

Why this answer

The MLflow Tracking UI is designed for experiment comparison. It presents runs in a sortable table with metrics and parameters, supports filtering, and allows drilling into individual runs. The Model Registry UI, Projects, and Jobs run history serve different purposes and do not offer the same run comparison and metric sorting features.

Exam trap

The trap here is assuming the Model Registry UI or Jobs run history can compare experiment runs, when only the MLflow Tracking UI provides that side-by-side metric view.

64
MCQmedium

An ML engineer trains a scikit-learn model on a Spark DataFrame in a Databricks notebook using MLflow autologging. The run logs parameters and metrics, but the engineer later opens the MLflow run and cannot find any input dataset lineage. Which action should the engineer take to ensure the training dataset is recorded with the run in the MLflow UI?

A.Set the MLFLOW_TRACKING_URI environment variable to the Unity Catalog metastore and rerun training.
B.Call mlflow.log_artifact() on the Spark DataFrame object directly after training.
C.Enable Delta table time travel on the source table and reference the table path in the run tags.
D.Wrap the training data with mlflow.data.from_spark() and pass it to mlflow.log_input().
AnswerD

mlflow.data.from_spark() constructs a Dataset object from a Spark DataFrame, and mlflow.log_input() records it on the active run, which is exactly what populates the Datasets section of the MLflow run UI. This is the supported mechanism for capturing input data lineage in MLflow on Databricks, and it works alongside autologging rather than replacing it.

Why this answer

MLflow records input dataset lineage only when a Dataset object is created and logged to the active run. The mlflow.data.from_spark() constructor builds that object from a Spark DataFrame, and mlflow.log_input() attaches it, after which the run UI shows dataset source, digest, and schema. Autologging captures params, metrics, and the model, but not input data lineage, so an explicit call is required.

Exam trap

The trap here is assuming MLflow autologging automatically captures input dataset lineage, when in fact dataset tracking must be added explicitly with mlflow.log_input().

65
MCQhard

Refer to the exhibit. A data scientist is attempting to log a model to an S3 bucket via MLflow, but they receive the error shown. What is the most likely root cause?

A.The S3 bucket is full and cannot accept new model artifacts.
B.The MLflow model flavor is not compatible with the S3 storage protocol.
C.The compute cluster's instance profile lacks write access to the S3 bucket.
D.The run_id 'xyz123' already exists and is locked by another user.
AnswerC

This is the most common cause for permission errors when interacting with S3 from Databricks. The instance profile or service principal attached to the cluster must have explicit IAM permissions granted to write to the specific bucket path. Without these permissions, all write operations will be blocked by the cloud provider.

Why this answer

The 'Permission denied' error indicates that the identity running the Databricks notebook or job lacks the necessary IAM permissions to write to the specified S3 bucket. In Databricks, interacting with external storage requires proper instance profile or service principal configuration. This error underscores the importance of cloud identity and access management (IAM) in ML workflows, as secure, authorized access to storage is foundational for maintaining the integrity and security of model artifacts during the development lifecycle.

Exam trap

Candidates often blame incorrect MLflow syntax for bucket write failures, overlooking cloud-level permissions like the compute cluster's instance profile configuration.

66
MCQhard

Refer to the exhibit. A Databricks job failed to start, returning the error shown. The job depends on MLflow for tracking. What is the most likely cause of this failure?

A.The Databricks Runtime version is too recent and contains a breaking change.
B.The cluster library configuration is missing the required MLflow package.
C.The MLflow tracking server is currently unreachable due to network security policies.
D.The job cluster has insufficient memory to load the MLflow dependency.
AnswerB

This error occurs when the driver or worker nodes lack the MLflow library. Since the job fails during initialization, it indicates that the environment definition for the job cluster does not include MLflow, preventing the code from accessing the tracking client at runtime during the start-up sequence.

Why this answer

The ClassNotFoundException indicates that the MLflow library is missing from the runtime environment of the job cluster. In Databricks, if a job is configured to use a custom or minimal environment, essential libraries must be explicitly installed. This error highlights the importance of dependency management in production ML workflows, ensuring that all necessary packages are present in the cluster environment to prevent execution failures during automated deployment or batch scoring jobs.

Exam trap

Test-takers frequently suspect corrupted model code when a job fails with a ClassNotFoundException, missing that required cluster libraries are simply uninstalled.

67
MCQmedium

A data scientist is using Databricks AutoML to train a classification model on a dataset with a binary target. They want to understand which features contributed most to the model's predictions and need a human-readable summary. Which AutoML output should they examine?

A.The MLflow run's artifact directory, specifically the 'feature_importance.json' file.
B.The AutoML-generated notebook that includes a section on model interpretability with SHAP values.
C.The Databricks Feature Store UI, which shows feature lineage and importance scores.
D.The Model Registry's version description, which is automatically populated with feature importance by AutoML.
AnswerB

Databricks AutoML generates a notebook for each trial that includes data exploration, model training, and interpretability using SHAP. The notebook provides summary plots and feature importance rankings, which are human-readable and explain how features affect predictions.

Why this answer

AutoML generates a notebook for each trial that includes model interpretability using SHAP. This notebook contains summary plots and feature importance rankings that are human-readable and explain the model's predictions. Examining this notebook is the correct way to understand feature contributions.

Exam trap

The trap here is assuming that feature importance is stored in a generic artifact or the Model Registry, when AutoML provides it in the generated interpretability notebook.

68
MCQhard

An ML engineer registers a model to the Databricks Model Registry and moves it to the 'Production' stage. A downstream batch scoring job references the model as models:/churn_model/Production. A data scientist then registers a new model version and transitions it to 'Production'. What happens to the downstream batch scoring job the next time it runs?

A.The job fails because a model URI cannot reference a stage once more than one version has occupied that stage.
B.The job continues to load the previously promoted version, because model URIs are cached and pinned at the time the job was first created.
C.The job loads the newly promoted Production version automatically, because the models:/churn_model/Production URI resolves to whichever version currently holds that stage.
D.The job loads the newest registered version regardless of stage, because the models:/ URI ignores stage names and always selects the highest version number.
AnswerC

Stage-based URIs such as models:/churn_model/Production are resolved at load time to the model version currently assigned to that stage. After the new version is transitioned to Production, the next job run picks up that version, enabling controlled rollout without changing job code, which is exactly the intended behavior.

Why this answer

Stage-based model URIs like models:/churn_model/Production resolve dynamically at load time to the version currently assigned to that stage. When a new version is transitioned into Production, subsequent scoring runs automatically use it, allowing controlled promotion without editing job code. The URI is not pinned or cached at job creation.

Exam trap

The trap here is assuming stage URIs are pinned when a job is defined, when they are actually resolved dynamically at each model load.

69
MCQhard

An ML engineer runs a hyperparameter tuning job on Databricks using Hyperopt with SparkTrials. The objective function trains a model and returns a validation metric. The engineer notices that each trial logs to the same MLflow run, making it impossible to compare trials. What should the engineer do to ensure each trial appears as a separate MLflow run?

A.Use mlflow.set_tag to assign a unique trial ID tag before logging metrics.
B.Pass a unique run_name argument to mlflow.log_metric for each trial.
C.Call mlflow.start_run(nested=True) at the beginning of the objective function.
D.Set the environment variable MLFLOW_EXPERIMENT_ID to a unique value inside the objective function.
AnswerC

SparkTrials creates a parent run for the tuning job and, when the objective function opens a nested run with mlflow.start_run(nested=True), each trial is logged as a child run under that parent. This yields one run per trial with its own parameters and metrics, enabling comparison in the MLflow UI. Databricks documents this pattern for distributed hyperparameter tuning.

Why this answer

With Hyperopt and SparkTrials on Databricks, the tuning job creates a parent MLflow run. To capture each trial separately, the objective function must open a nested run using mlflow.start_run(nested=True). This produces child runs per trial, each with its own parameters and metrics, which can be compared in the MLflow experiment UI.

Tags, metric names, or experiment changes do not create distinct runs.

Exam trap

The trap here is thinking that tagging or renaming metrics separates trials, when only creating a nested run per trial actually produces distinct MLflow runs.

70
MCQmedium

A data scientist trains a scikit-learn model in a Databricks notebook and calls mlflow.sklearn.log_model with the registered_model_name argument set to 'churn_model'. Later, a colleague needs to know which source notebook and Git commit produced run ID 3f8a1c. Where can this information be retrieved?

A.In the MLflow run's artifact directory, open the conda.yaml file, which records the source notebook path and Git commit used during training.
B.In the Databricks Model Registry, open version 1 of churn_model and read the 'Source' field, which stores the originating notebook path and Git commit.
C.In the MLflow Experiment page, open the run and inspect the run's tags, including mlflow.source.name and mlflow.source.git.commit.
D.In the Databricks Jobs UI, locate the task that ran the notebook and read its run history, which stores the MLflow run ID and its source metadata.
AnswerC

MLflow automatically records the source notebook path, Git commit, and other environment metadata as tags on every run. Opening the run in the experiment UI and viewing its tags reveals exactly which notebook and Git revision generated run 3f8a1c, satisfying the colleague's need without any extra instrumentation.

Why this answer

MLflow automatically attaches metadata tags to each run, including the source notebook path, Git commit, and user. These tags are visible on the run detail page in the experiment UI, making it the authoritative place to trace a run back to its originating notebook and revision. Registry entries and artifact files like conda.yaml do not carry this provenance.

Exam trap

The trap here is assuming the Model Registry stores full source provenance, when that metadata actually belongs to the MLflow run tags in the experiment.

71
MCQmedium

An ML engineer wants to package a training project so it can be run reproducibly from a Databricks Job across environments, with its Python dependencies and entry point defined. Which MLflow capability should be used?

A.MLflow Tracking, by logging parameters and metrics for each run.
B.MLflow Recipes, by defining a step-based pipeline with profiles.
C.MLflow Projects, using an MLproject file to declare the entry point and environment.
D.MLflow Models, by saving the estimator with a signature.
AnswerC

MLflow Projects define a reusable, reproducible unit through an MLproject file that names the entry point, parameters, and environment specification such as a conda or pip requirements file. Running the project creates a run with recorded parameters and source, which supports reproducibility across environments and integrates with Databricks Jobs as a project task.

Why this answer

MLflow Projects are the packaging mechanism for reproducible runs: the MLproject file specifies the entry point and the environment via a conda or pip specification, and executing the project records source and parameters in a run. Models capture artifacts, Tracking records results, and Recipes impose a specific template, so none of those define a reusable training entry point with dependencies for a Databricks Job.

Exam trap

The trap here is conflating MLflow Models or Tracking with reproducibility packaging, when defining an entry point and dependencies is the job of MLflow Projects.

72
Multi-Selecthard

Which THREE actions are best practice when deploying a machine learning model using Databricks Model Serving?

Select 3 answers
A.Hardcode API credentials directly into the model inference script.
B.Ensure the model signature is defined to enable input validation.
C.Utilize the Model Registry to manage the versioning of the deployed model.
D.Perform inference on the same cluster used for training to save costs.
E.Implement logging within the inference function to monitor performance.
AnswersB, C, E

Defining a model signature allows the serving endpoint to validate incoming request data against the expected schema. This prevents runtime errors and unexpected model behavior by rejecting malformed input, which is a crucial safeguard for stable and reliable production-grade ML inference services in a distributed environment.

Why this answer

Databricks Model Serving provides a managed endpoint for low-latency inference. Adopting best practices for deployment—such as using environment variables, ensuring proper logging, and testing in staging—is essential for maintaining reliable, scalable, and secure production services. These practices ensure that the serving infrastructure is decoupled from the training environment, adheres to security policies, and provides observability into model performance in real-world conditions, effectively bridging the gap between model development and operational deployment.

Exam trap

Candidates frequently overlook model signatures as optional, whereas they are mandatory best practices for payload validation and seamless model serving integration.

73
MCQhard

An ML engineer has registered a model in the Databricks Model Registry. The model must be deployed to a REST endpoint that automatically scales with traffic and provides a stable serving environment. Which Databricks capability should they use?

A.Export the model as an MLflow artifact and load it in a Databricks notebook using mlflow.pyfunc.load_model for interactive scoring.
B.Databricks Model Serving with the registered model version as the served entity.
C.MLflow Model Registry webhooks that trigger a deployment script when the model version transitions to Production.
D.A Databricks job that runs a Python web server on a cluster and exposes the model via a public URL.
AnswerB

Databricks Model Serving is a fully managed, serverless solution that deploys registered model versions as REST endpoints with automatic scaling and a stable serving environment. It handles infrastructure, scaling, and monitoring, so the engineer only needs to select the model version and enable serving, meeting the deployment requirement without managing servers.

Why this answer

Databricks Model Serving provides a managed, serverless endpoint for registered model versions. It automatically scales with request volume and offers a stable serving environment, eliminating the need to run and maintain a web server. The other options either require manual server management or are intended for interactive or batch use, not production REST serving.

Exam trap

The trap here is thinking that any Python web server on a cluster constitutes model serving, when only Databricks Model Serving provides managed autoscaling and a stable endpoint.

74
MCQmedium

A data scientist wants to package a training script so that it can be run repeatedly with different hyperparameters and shared with colleagues who use different Python library versions. The script must create a reproducible environment. Which MLflow component should they use to define the project and its dependencies?

A.MLflow Projects with an MLproject file and a conda.yaml environment specification.
B.MLflow Tracking with a dedicated experiment and run tags for each library version.
C.MLflow Models with a custom Python function flavor and a requirements.txt file.
D.MLflow Model Registry with stage transitions and model version descriptions.
AnswerA

MLflow Projects provide a standard format for packaging reusable code. The MLproject file defines entry points and parameters, while conda.yaml specifies the exact library versions. This allows colleagues to run the project with mlflow run and get a reproducible environment, satisfying both the parameterization and dependency isolation requirements.

Why this answer

MLflow Projects are designed to package code with a clear entry point and parameter definitions, and the conda.yaml file locks dependency versions for reproducibility. This lets colleagues run the same training script with different hyperparameters in an isolated environment. Tracking, Models, and the Model Registry serve different lifecycle stages and do not provide project packaging.

Exam trap

The trap here is equating MLflow Models or the Model Registry with project packaging, when only MLflow Projects define runnable entry points and environment specifications.

75
MCQmedium

You are using Databricks Feature Store to create a feature table that will be used for both batch training and online inference. The feature table must be refreshed daily with new data, and the online store must serve the latest feature values within minutes of the refresh. Which configuration should you use?

A.Create a feature table with online=False, and configure Model Serving to read from the offline table at inference time.
B.Create a standard Delta table and manually export it to a key-value store each day using a notebook.
C.Create a feature table with online=True and use a Databricks Job to run a daily write that updates both the offline Delta table and the online store.
D.Create a feature table with online=True, but only write to it weekly to reduce costs.
AnswerC

Setting online=True when creating the feature table enables an online store (such as DynamoDB or SQL) alongside the offline Delta table. A daily job that writes to the feature table via FeatureStoreClient.write_table updates both stores, and the online store is refreshed with the new values. This meets the requirement for daily refresh and near-real-time online serving, as Databricks Feature Store publishes to the online store during the write.

Why this answer

A feature table created with online=True is backed by both an offline Delta table and an online store. Writing to it via FeatureStoreClient.write_table publishes the new feature values to both stores. Running this write in a daily Databricks Job refreshes the online store with the latest data, enabling low-latency serving within minutes.

Other options either omit the online store, use unsupported manual exports, or fail to meet freshness requirements.

Exam trap

The trap here is assuming that Model Serving can read directly from the offline Delta table, when online inference requires a feature table published to an online store.

Page 1 of 2 · 83 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Ml Workflows questions.