MLA-C01 Data Preparation for Machine Learning Practice Question
A machine learning engineer is preparing data for a SageMaker training job and needs to split a large dataset into training, validation, and test sets while avoiding data leakage from the same entity appearing in multiple splits. The dataset contains multiple rows per customer, and the target is customer churn. Which TWO strategies are appropriate? (Choose two.)
⚠ Common exam trap
Test-takers frequently confuse reproducibility with leakage prevention, since a fixed random seed produces repeatable splits but still scatters a customer's rows across training and validation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the data by a hash of the customer ID so that all rows for a given customer land in exactly one split.
Entity-level leakage occurs when rows from the same customer appear in more than one split, letting the model memorize customer-specific patterns and inflating validation metrics. Hashing the customer ID into split buckets and using GroupShuffleSplit with the customer ID as the group both ensure that all rows for a customer stay together, which is the correct way to partition this churn dataset.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Split the data by a hash of the customer ID so that all rows for a given customer land in exactly one split.
Why this is correct
Hashing the customer ID and assigning splits by hash range guarantees that every row for a customer goes to the same split, which prevents the same entity from appearing in both training and validation. This is a standard grouped-split technique for entity-level leakage and is deterministic and reproducible across reruns when the same hash function and boundaries are used.
- ✗
Duplicate rows for customers in the minority class so that each split contains examples of every customer.
Why it's wrong here
Duplicating rows so that customers appear in every split is the opposite of the required behavior: it deliberately places the same entity in multiple splits and will cause severe leakage. While oversampling can address class imbalance, it must be applied within the training split only, after grouping, and never in a way that spreads a customer's data across validation or test sets.
- ✗
Sort the dataset by timestamp and take the first 80 percent of rows for training and the last 20 percent for validation.
Why it's wrong here
A chronological split is appropriate when the goal is to simulate future data, but it does not guarantee that a customer's rows are confined to one split. If customers recur over time, the same entity can appear in both the training and validation windows, which reintroduces leakage. This strategy addresses temporal ordering, not entity grouping, so it does not satisfy the requirement.
- ✓
Use GroupShuffleSplit from scikit-learn with the customer ID as the group parameter to partition customers across splits.
Why this is correct
GroupShuffleSplit assigns entire groups, defined by the customer ID, to a single split, which directly enforces entity-level separation between training, validation, and test sets. It is the purpose-built scikit-learn utility for this scenario and can be run inside a SageMaker Processing job to produce the split datasets before training begins.
- ✗
Use a random row-level split with a fixed seed, because a fixed seed makes the split reproducible and therefore leakage-free.
Why it's wrong here
A fixed seed makes a random split reproducible, but it does not prevent the same customer's rows from being distributed across training and validation. Reproducibility and leakage prevention are different properties. With multiple rows per customer, a row-level random split will almost certainly place some of a customer's rows in each split, inflating validation performance and misleading model selection.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.