Courseiva
Model Development →mediumMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

A data scientist is using Databricks Feature Store to build a training set for a model that predicts customer churn. The feature table contains a column `customer_id` and several features, and the label is stored in a separate Delta table. The data scientist wants to ensure that the exact same feature values used during training are available at inference time. Which approach correctly uses Databricks Feature Store to create the training set?

⚠ Common exam trap

Test-takers frequently confuse the training set creation with model logging; `log_model` is for logging models that will use Feature Store features at inference, not for creating the training set.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use `fs.create_training_set()` with the feature table and label DataFrame, specifying `customer_id` as the lookup key.

The `create_training_set` method is specifically designed to join feature tables with labels using a lookup key, preserving feature lineage and ensuring that the same feature computations are used during training and inference. This prevents training-serving skew and simplifies deployment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use `fs.create_training_set()` with the feature table and label DataFrame, specifying `customer_id` as the lookup key.

    Why this is correct

    The `create_training_set` method of the FeatureStoreClient joins feature tables with labels using the specified lookup key, producing a training set that includes the feature values and lineage. This ensures consistency because the same feature computation is used for both training and inference.

  • ✗

    Perform a manual join between the feature table and the label table using Spark, then log the resulting DataFrame as an MLflow artifact.

    Why it's wrong here

    A manual join does not capture the feature lineage or ensure that the same feature values are used at inference. Databricks Feature Store is designed to manage this consistency; bypassing it with a manual join loses the automatic tracking and may lead to training-serving skew.

  • ✗

    Export the feature table to a CSV file and merge it with the label data using pandas, then log the model with MLflow.

    Why it's wrong here

    Exporting to CSV and using pandas breaks the connection to the Feature Store. At inference, the model would not automatically look up features from the online store, and the manual merge process is error-prone and not scalable for large datasets.

  • ✗

    Use `fs.log_model()` with the feature table name and specify the label column; the method automatically creates the training set.

    Why it's wrong here

    `log_model` is used to log a model that uses features from the Feature Store, but it does not create the training set. The training set must be created separately using `create_training_set`, and then the model can be logged with `log_model` to ensure feature lookup at inference.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.