Courseiva

MLA-C01 Data Preparation for Machine Learning Practice Question

A data scientist is preparing a dataset for a machine learning model that predicts customer churn. The dataset contains a column 'CustomerID' that is a unique identifier. What should the data scientist do with this column before training the model?

⚠ Common exam trap

Test-takers frequently think unique identifiers are useful for tracking or that they can be encoded as categorical features, but the exam tests the principle that identifiers with no predictive relationship to the target must be removed to avoid overfitting and data leakage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Remove the column from the feature set.

'CustomerID' is a unique identifier with no predictive power for churn. Including it as a feature would cause the model to memorize individual customers rather than learn generalizable patterns, leading to overfitting and poor performance on unseen data. In machine learning, such columns should be removed during data preparation to ensure the model learns from meaningful features.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Keep the column as a feature because it uniquely identifies each customer.

    Why it's wrong here

    A unique identifier carries no generalisable signal: each value appears once, so the model memorises training rows rather than learning churn patterns, and unseen customers cannot be scored. Dropping the column is the correct preprocessing step. Keeping identifiers is legitimate when the task is record lookup or entity resolution, where per-row identity is the prediction target itself.

  • ✗

    Use the column as the target variable.

    Why it's wrong here

    CustomerID is an arbitrary label with no relationship to churn, so a model trained to predict it learns nothing about customer behaviour and cannot generalise. The target must be the churn indicator. Identifiers serve as targets only in tasks such as record linkage or entity resolution, where matching an identity is the objective.

  • ✓

    Remove the column from the feature set.

    Why this is correct

    CustomerID is a unique identifier carrying no generalisable signal; each value appears once, so the model would memorise rather than learn. Removing it from the feature set prevents noise and spurious splits, satisfying the requirement to prepare the churn dataset correctly.

  • ✗

    Encode the column using one-hot encoding.

    Why it's wrong here

    One-hot encoding a unique identifier creates one binary column per customer, producing a sparse matrix with as many features as rows; tree and linear models gain no predictive signal and dimensionality explodes. The column should be dropped. One-hot encoding suits low-cardinality categorical variables such as contract type or payment method, where each category recurs across many customers.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.