Courseiva

CCNA Ml Workflows Questions

8 of 83 questions · Page 2/2 · Ml Workflows topic · Answers revealed

76
MCQhard

You are deploying a model using MLflow, and you want to log the custom pre-processing logic alongside the model so it is automatically applied during inference. How should you achieve this?

A.Create a separate inference job that runs the pre-processing code before the model.
B.Use the mlflow.pyfunc.log_model() function with a custom code wrapper.
C.Store the pre-processing logic in a shared Delta table and call it at inference.
D.Log the pre-processing parameters as tags in the MLflow run and re-apply them.
AnswerB

The pyfunc flavor is the standard way to package custom logic with models in MLflow. It allows users to define a wrapper that executes arbitrary code, including pre-processing, at inference time. This ensures that the model is self-contained, simplifying deployment and guaranteeing that input data is processed correctly.

Why this answer

MLflow allows you to log custom Python functions or entire pipelines as part of the model using the 'pyfunc' flavor. By wrapping your pre-processing steps and model into a custom pyfunc class and logging it, you encapsulate both the transformation and the prediction logic. This is a critical best practice for production deployments, as it ensures consistent pre-processing between training and inference, eliminating training-serving skew and simplifying the deployment process for the serving endpoint.

Exam trap

Test-takers often select generic artifact logging functions instead of the specialized pyfunc wrapper required to bundle custom pre-processing logic with the model.

77
MCQmedium

A data scientist needs to track parameters, metrics, and model artifacts during training on Databricks. Which approach is the industry-standard best practice to ensure reproducibility and lineage?

A.Manually save model artifacts to DBFS and document parameters in a spreadsheet.
B.Write all model logs to a local temporary folder on the driver node.
C.Use the MLflow Tracking API to log parameters and artifacts during the training run.
D.Deploy the model directly to production without logging parameters to save compute costs.
AnswerC

MLflow Tracking provides a robust, centralized API to record experiments. It automatically captures metadata, code versions, and artifacts, making them accessible via the Experiments sidebar. This ensures that every model training process is documented, searchable, and repeatable, which is foundational for maintaining high-quality machine learning model lifecycle management.

Why this answer

MLflow Tracking is the core component for logging parameters, metrics, and artifacts in Databricks. By integrating MLflow into the training script, data scientists create a centralized record of every experiment run. This allows for easy comparison between different model versions, facilitates team collaboration, and ensures that the exact code, environment, and hyperparameter configuration used to produce a specific model artifact can be audited and reproduced later in the MLOps lifecycle.

Exam trap

Candidates often rely on manual logging or external databases, ignoring the built-in MLflow Tracking API which is the industry standard for maintaining lineage and reproducibility in Databricks.

78
MCQmedium

You are building a pipeline in Databricks and need to ensure that a training job only runs after the upstream data preparation job has successfully completed. Which Databricks feature should you use?

A.Use the mlflow.end_run() command at the end of the data prep script.
B.Use Delta Live Tables (DLT) expectations to trigger the job.
C.Use Databricks Workflows task dependencies.
D.Use a cron schedule with a 30-minute delay for the training job.
AnswerC

Databricks Workflows enables users to build multi-task jobs where tasks are linked by success dependencies. This is the native and most efficient way to ensure that a training task waits for the data preparation task to complete, maintaining a strict and reliable execution order within the pipeline.

Why this answer

Databricks Workflows allows for the creation of complex pipelines using task dependencies. By defining a dependency between the data preparation task and the training task, the orchestrator ensures the correct execution order. This dependency management is foundational for building reliable end-to-end ML pipelines, ensuring that models are always trained on processed, valid data and preventing failures caused by missing or stale input data during the automated training process.

Exam trap

Candidates often try to manage dependencies using custom notebook code or external scheduling tools. They fail to recognize that Databricks Workflows has built-in native support for task dependencies.

79
MCQmedium

Refer to the exhibit. A machine learning engineer wants to promote this model to Production. Which MLflow action should be performed to achieve this while ensuring the model is ready for deployment?

A.Use the mlflow.register_model() function with the production flag.
B.Call the transition_model_version_stage API to move the model to Production.
C.Delete the model from Staging and re-upload it to the Production path.
D.Update the tag 'status' to 'Production' in the model metadata.
AnswerB

The transition_model_version_stage API is the standard method for updating the stage of a registered model version. By programmatically changing the stage to 'Production', the model becomes available for production deployment services to fetch and serve the most current approved model artifact for live traffic.

Why this answer

The MLflow Model Registry provides the `transition_model_version_stage` function to promote models between stages like Staging and Production. Promoting a model involves updating its state in the registry, which then allows downstream deployment pipelines to identify the correct model version. This structured approach ensures that only verified models are promoted to production environments, maintaining governance and stability in the model deployment process.

Exam trap

Candidates often confuse manual model logging with registry stage transitions, incorrectly selecting options related to saving files or creating new experiments instead of using the specific API for stage promotion.

80
MCQhard

An ML engineer is configuring a Databricks Job to retrain a model daily. The job must first run a notebook that creates a feature table, then run a notebook that trains and registers the model. The engineer wants to ensure the training notebook only runs if the feature table creation succeeds. Which Databricks Workflows feature should they use to define this dependency?

A.Set the training task's depends_on field to the feature table task.
B.Set the training task's run_if condition to ALL_SUCCESS, which automatically includes the feature table task.
C.Use a conditional task with a run_if condition based on the feature table task's output.
D.Configure the feature table task to call the training notebook at the end of its run.
AnswerA

In Databricks Workflows, each task can specify a list of upstream tasks in its depends_on field. Setting the training task's depends_on to the feature table task ensures the training task runs only after the feature table task completes successfully. This creates a directed acyclic graph (DAG) that enforces the required order and success condition.

Why this answer

In Databricks Workflows, task dependencies are defined using the depends_on field, which lists upstream tasks that must complete successfully before the task runs. This creates a DAG and ensures the training task executes only after the feature table task succeeds. Other options either do not establish the dependency or rely on additional configuration that still requires depends_on.

Exam trap

The trap here is thinking that run_if conditions alone can establish dependencies without explicitly listing upstream tasks in depends_on.

81
MCQhard

An ML engineer is using Databricks Jobs to orchestrate a pipeline that includes a notebook for feature engineering and a notebook for model training. The training notebook must run only after the feature engineering notebook completes successfully, and both must run on a schedule. Which configuration in Databricks Jobs achieves this dependency?

A.Use a single notebook that calls the feature engineering notebook and then the training notebook using %run.
B.Set the training notebook to run on a schedule that starts after the expected completion time of the feature engineering notebook.
C.Create a multi-task job where the training task has a depends_on condition referencing the feature engineering task.
D.Create two separate jobs and configure the training job to be triggered by the completion of the feature engineering job using a job run trigger.
AnswerC

Databricks Jobs supports multi-task workflows with dependencies. By setting depends_on for the training task to the feature engineering task, the training task runs only after the feature engineering task succeeds. This is the native, supported way to express task dependencies within a job, and the job can be scheduled.

Why this answer

Databricks Jobs multi-task workflows allow defining tasks with dependencies using depends_on. This ensures the training task starts only after the feature engineering task succeeds, and the entire job can be scheduled. Separate jobs with triggers, %run, or time-based scheduling do not provide the same reliable dependency semantics.

Exam trap

The trap here is thinking that separate jobs or time-based scheduling can enforce a dependency, when only a multi-task job with depends_on guarantees sequential execution on success.

82
MCQhard

An ML engineer is using Databricks Jobs to orchestrate a multi-step ML pipeline. The pipeline includes a task that trains a model and logs it to MLflow, followed by a task that registers the model in the Model Registry. The engineer wants to ensure that the model is only registered if its accuracy exceeds a threshold. Which approach best implements this conditional logic?

A.Configure the Model Registry to automatically register models only if they meet a performance threshold set in the model's metadata.
B.In the training task, evaluate the model and if accuracy exceeds the threshold, call the MLflow Model Registry API to register the model.
C.Use a Databricks Job with a conditional task that checks a metric value stored in a Delta table and then triggers registration.
D.Use a Databricks Job with multiple tasks and define a dependency that runs the registration task only if the training task succeeds.
AnswerB

Embedding the evaluation and conditional registration within the training task allows the logic to check accuracy and only register if the threshold is met. This is a common pattern in Databricks Jobs, where a single task can contain multiple steps. It ensures that registration is gated by the metric, avoiding unnecessary separate tasks.

Why this answer

The most straightforward way to conditionally register a model based on accuracy is to perform the evaluation within the training task and call the Model Registry API only if the threshold is met. This keeps the logic in one place and avoids unnecessary task dependencies. Other approaches either do not enforce the metric condition or require additional complexity.

Exam trap

The trap here is assuming that task dependencies or Model Registry settings can enforce metric-based conditions, but they only handle success/failure or metadata, not dynamic thresholds.

83
MCQeasy

An ML engineer is using MLflow Tracking to compare multiple runs of a hyperparameter tuning experiment. The engineer wants to quickly identify the run that achieved the best validation accuracy and then promote that run's model to the Model Registry. Which MLflow feature allows the engineer to view and compare runs in a centralized UI?

A.MLflow Model Registry
B.MLflow Projects
C.MLflow Tracking UI
D.MLflow Models
AnswerC

The MLflow Tracking UI provides a centralized interface to view experiments, runs, metrics, parameters, and artifacts. It allows sorting and filtering runs by metrics, making it easy to identify the best run. From the UI, you can also register a model directly to the Model Registry. This directly addresses the engineer's need to compare and promote.

Why this answer

The MLflow Tracking UI is designed to log and compare runs, displaying metrics, parameters, and artifacts in a sortable table. It allows the engineer to identify the best run and then register its model to the Model Registry with a few clicks, streamlining the workflow.

Exam trap

The trap here is confusing the Model Registry with the Tracking UI; the Registry manages model versions, while the Tracking UI compares runs.

← PreviousPage 2 of 2 · 83 questions total

Ready to test yourself?

Try a timed practice session using only Ml Workflows questions.