AIF-C01 Fundamentals of AI and ML Practice Question
A startup is preparing a dataset to train a supervised machine learning model and wants to follow sound data preparation practices. Which TWO activities are appropriate parts of preparing training data? (Choose two.)
⚠ Common exam trap
The trap here is treating 'cleaner and bigger' as always better, so deleting every incomplete row or duplicating records looks reasonable even though both harm data quality or evaluation validity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Splitting the dataset into training, validation, and test subsets before tuning the model.
Sound preparation centers on protecting evaluation integrity and improving data quality: separating train, validation, and test subsets keeps performance estimates honest, while removing duplicates and fixing inconsistent labels prevents leakage and contradictory signals. Tuning on the test set, deleting all incomplete rows, and duplicating examples either leak information, discard useful data, or add no new information.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Splitting the dataset into training, validation, and test subsets before tuning the model.
Why this is correct
Holding out validation and test subsets lets the team tune hyperparameters on validation data and estimate final performance on untouched test data. Without this separation, performance estimates become optimistically biased because the model has effectively seen the evaluation data during development, so this is a core preparation practice.
- ✗
Adding synthetic copies of the same training examples until the dataset reaches a target size.
Why it's wrong here
Duplicating existing examples does not add new information and can cause the model to overweight those records, effectively worsening overfitting. If more data is needed, augmentation or collection of genuinely new examples is preferable, so simple duplication is not sound preparation.
- ✓
Removing duplicate records and correcting inconsistent label values across the dataset.
Why this is correct
Duplicates can cause the same example to appear in both training and evaluation splits, inflating measured performance, while inconsistent labels teach the model contradictory patterns. Cleaning these issues during preparation improves data quality and makes reported metrics more trustworthy, so it is a valid preparation activity.
- ✗
Tuning the model's hyperparameters repeatedly against the test set until the score improves.
Why it's wrong here
Repeatedly tuning against the test set leaks information from the evaluation data into model selection, producing an overly optimistic final estimate that will not hold in production. The test set must remain untouched until the very end, so this is a misuse of data rather than a preparation practice.
- ✗
Deleting all records that contain any missing values to guarantee a clean dataset.
Why it's wrong here
Dropping every row with a missing value can discard a large share of data and introduce bias if missingness correlates with the target. Imputation or explicit missing-value handling often preserves more information, so blanket deletion is not a recommended preparation practice.
Go deeper
Related to this question
About these practice questions
One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.