Courseiva

CCNA Ml Assoc Model Development Questions

10 of 85 questions · Page 2/2 · Ml Assoc Model Development topic · Answers revealed

76
MCQhard

Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?

A.It forces the model to run on a GPU during inference.
B.It speeds up the model training process.
C.It enables schema validation and improves model serving usability.
D.It reduces the storage size of the model artifact.
AnswerC

Signatures provide a clear contract for the model, enabling automatic validation of input data. Input examples help serving tools generate accurate API definitions, making it easier for external applications to call the model. This reduces integration friction and ensures that only valid data is passed to the model.

Why this answer

Providing a signature and input example allows MLflow to define the expected schema for the model. This metadata facilitates automatic type validation and allows the model serving infrastructure to generate accurate API documentation. This is critical for downstream consumers who need to integrate the model, as it prevents runtime integration errors and clarifies the interface expectations before the model is even deployed into production.

Exam trap

Candidates often view signatures as optional metadata, failing to realize they provide the essential interface documentation and validation required for seamless downstream integration by other services.

77
MCQeasy

A data scientist is training a linear regression model using scikit-learn on Databricks. They want to track the model's hyperparameters, such as fit_intercept and normalize, in MLflow. Which MLflow API call should they use to log these hyperparameters?

A.mlflow.log_artifact
B.mlflow.log_param
C.mlflow.log_metric
D.mlflow.set_tag
AnswerB

mlflow.log_param logs a single key-value pair representing a hyperparameter. It is the correct API for recording settings like fit_intercept and normalize. Each call logs one parameter, and parameters are immutable once logged for a run, ensuring a clear record of the configuration used for that training run.

Why this answer

To track hyperparameters in MLflow, the appropriate API is mlflow.log_param, which records key-value pairs that define the model's configuration. This allows the parameters to be visible in the MLflow UI and searchable via the API. Logging metrics, artifacts, or tags would not correctly capture hyperparameters and would hinder experiment comparison.

Exam trap

The trap here is assuming that any logging function can record hyperparameters; only log_param is designed for that purpose.

78
MCQmedium

A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune a scikit-learn model. They set max_evals=100 and parallelism=4. After the tuning completes, they notice that some trials failed due to memory errors on the workers. What is the most likely cause of these failures?

A.The scikit-learn model is being trained on the driver, causing memory pressure on the driver instead of workers.
B.Each trial is training on the full dataset, and the dataset is too large to fit in the memory of a single worker.
C.SparkTrials does not support scikit-learn models and should be used only with MLlib.
D.The cluster's autoscaling is reducing the number of workers, causing trials to be queued and eventually fail.
AnswerB

SparkTrials runs each trial on a single worker, and if the dataset is large, it may exceed the worker's memory. Unlike distributed training, each trial is independent and uses the full dataset unless you subsample or use Spark ML. This leads to out-of-memory errors when the data is too big for one worker.

Why this answer

The correct answer is that each trial is training on the full dataset, and the dataset is too large to fit in the memory of a single worker. SparkTrials distributes trials, but each trial runs on one worker with the full dataset. If the data is large, it can cause out-of-memory errors.

To mitigate, you can reduce the dataset size, increase worker memory, or use distributed algorithms. The other options misattribute the cause to model incompatibility, driver memory, or autoscaling.

Exam trap

The trap here is assuming that SparkTrials distributes the training of a single model across workers, when it actually distributes independent trials, each on a single worker.

79
MCQeasy

A data scientist is working in a Databricks notebook and wants to view the results of their MLflow runs, including metrics and parameters, directly within the notebook. Which MLflow function should they use?

A.mlflow.search_runs()
B.mlflow.get_experiment_by_name()
C.mlflow.get_run()
D.mlflow.list_experiments()
AnswerA

mlflow.search_runs() returns a pandas DataFrame containing run information such as parameters and metrics from the specified experiment. It can be used within a notebook to programmatically inspect and analyze runs. This function is ideal for displaying results directly in the notebook environment, as it provides a structured view of all runs.

Why this answer

mlflow.search_runs() is designed to query runs within an experiment and return a pandas DataFrame with columns for run ID, parameters, metrics, and other metadata. This makes it straightforward to display and analyze multiple runs in a notebook. Other functions either retrieve single runs or list experiments without providing run-level details in a tabular format.

Exam trap

The trap here is confusing functions that list experiments or retrieve single runs with the one that returns a comprehensive table of runs for analysis.

80
MCQmedium

A machine learning engineer is using MLflow to log a model built with XGBoost. They call mlflow.xgboost.log_model(xgb_model, 'model') and then attempt to load the model in a different environment using mlflow.pyfunc.load_model('runs:/<run_id>/model'). The load fails with an error about missing dependencies. Which action should they take to ensure the model can be loaded in the new environment?

A.They should log the model with the registered_model_name parameter to ensure dependencies are captured.
B.They should save the model in the ONNX format instead, because ONNX models have no dependencies.
C.They should include the XGBoost library in the conda environment when logging the model, either by passing a custom conda_env or by ensuring the library is installed in the current environment so MLflow can infer it.
D.They should use mlflow.xgboost.load_model() instead of mlflow.pyfunc.load_model(), because the latter does not support XGBoost models.
AnswerC

This is correct because MLflow captures dependencies in a conda.yaml file when logging a model. If XGBoost is not included, loading in a new environment will fail. By default, MLflow infers dependencies from the current environment, but if XGBoost was not installed or not detected, it may be missing. Providing a custom conda_env or ensuring XGBoost is installed during logging ensures the dependency is recorded.

Why this answer

When logging a model with MLflow, dependencies are recorded in a conda environment file. If a required library like XGBoost is missing from that file, loading in a new environment fails. Ensuring the library is installed during logging allows MLflow to capture it, or you can specify a custom conda environment.

This guarantees the model can be reproduced.

Exam trap

The trap here is assuming that model registration or using a native loader solves dependency issues, when the real fix is to include the dependency in the logged environment.

81
Multi-Selectmedium

Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?

Select 2 answers
A.Hardcode the data file paths to local desktop directories.
B.Log the Git commit hash along with experiment metadata in MLflow.
C.Only log the model artifact and ignore the environment dependencies.
D.Utilize Delta Time Travel to query the data as it existed during training.
E.Manually delete all previous MLflow runs to keep the tracking server clean.
AnswersB, D

Logging the Git commit hash allows developers to link the model artifact to the exact version of the source code that produced it. This is essential for auditing and reproducing experiments months later, as it provides a clear record of the code changes that led to specific model performance.

Why this answer

Reproducibility in Databricks is achieved by capturing the state of the code, the environment, and the data. By utilizing Git integration and MLflow tracking, data scientists can point to the exact version of the code and environment used for an experiment. Using versioned Delta tables ensures that the data state can also be reconstructed, preventing 'data drift' from invalidating the results of historical experiments as underlying data tables change.

Exam trap

Candidates often forget that code and environment versioning are insufficient without capturing the exact historical state of the underlying data.

82
MCQmedium

A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?

A.One-hot encode all categorical variables.
B.Use Target Encoding to map categories to their mean target value.
C.Drop the categorical features entirely.
D.Convert all features into a single string representation.
AnswerB

Target encoding is highly effective for high-cardinality features, as it transforms categorical levels into a continuous numerical representation based on the target variable. This reduces dimensionality while retaining meaningful information, which is ideal for linear models that perform best when features are appropriately scaled and numeric.

Why this answer

Linear models require numeric inputs and are sensitive to the scale and representation of categorical variables. One-hot encoding high-cardinality features results in sparse, high-dimensional matrices that degrade performance. Target encoding or using specialized embeddings is more efficient.

Choosing the right encoding strategy is a critical model development decision that directly impacts training speed, convergence, and the overall predictive accuracy of the final model deployed into production.

Exam trap

Candidates often default to One-Hot Encoding, ignoring that high-cardinality categorical features cause dimensionality explosion, making Target Encoding the more scalable and performant choice for linear models.

83
MCQmedium

Refer to the exhibit. Why is providing an 'input_example' highly recommended during the model logging process?

A.It automatically generates unit tests for the training pipeline.
B.It enables automatic schema validation and documentation in the model registry.
C.It increases the accuracy of the model during inference.
D.It forces the model to use a faster serializing algorithm.
AnswerB

The input example allows the system to verify that the model's expected inputs match the actual provided data. This is a key feature of Unity Catalog, which uses this example to document the model, making it discoverable and ensuring that consumers have a clear understanding of the expected input format.

Why this answer

Providing an input example allows MLflow to infer the model's input schema and provides a test case for downstream deployment. This example is critical for Databricks to validate the model's interface, allowing automated tests to run against the model immediately upon deployment. It also serves as a form of self-documentation, helping other team members understand how to interact with the model's API without needing to look at complex training code.

Exam trap

Candidates often view 'input_example' as optional documentation, failing to realize it is a functional requirement for MLflow to infer schemas and perform automated validation during deployment.

84
MCQhard

A machine learning engineer is using MLflow to log a model trained with XGBoost. They want to ensure that the model can be loaded and used for inference in a different environment without requiring the original training environment. Which MLflow feature allows the model to capture its dependencies and environment?

A.MLmodel file with conda_env and requirements.txt
B.Run ID
C.Model signature
D.Model version
AnswerA

When logging a model, MLflow automatically creates an MLmodel file that includes a conda environment specification and a requirements.txt file. These files list the dependencies needed to recreate the environment. This allows the model to be loaded in a different environment by installing the specified packages, ensuring reproducibility.

Why this answer

The MLmodel file, along with conda_env and requirements.txt, captures the dependencies required to load and run the model. When logging a model, MLflow automatically generates these files based on the current environment. This enables the model to be deployed in a different environment by recreating the necessary dependencies, ensuring consistent behavior.

Exam trap

The trap here is thinking that the model signature or version handles environment portability; only the MLmodel file with dependency specifications does that.

85
MCQeasy

A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?

A.df.randomSplit([0.8, 0.2], seed=42)
B.df.repartition(2)
C.df.sample(fraction=0.8, seed=42)
D.df.cache()
AnswerA

randomSplit splits a DataFrame into multiple DataFrames according to the provided weights, and the seed parameter ensures reproducibility. Using seed=42 guarantees that the same split is generated each time the code runs, which is essential for consistent model evaluation. This method is the standard PySpark approach for train-test splitting with a fixed random seed.

Why this answer

randomSplit is the PySpark method designed to split a DataFrame into multiple parts based on weights, and it accepts a seed for reproducibility. This allows the data scientist to obtain consistent training and testing sets across runs. Other methods like sample, repartition, or cache do not provide a reproducible train-test split.

Exam trap

The trap here is confusing sampling with splitting, when randomSplit is specifically built to return multiple disjoint DataFrames with a seed for reproducibility.

← PreviousPage 2 of 2 · 85 questions total

Ready to test yourself?

Try a timed practice session using only Ml Assoc Model Development questions.