Courseiva

AI0-001 Implementing AI Solutions Practice Question

A data scientist is preparing a dataset for a text classification model. To prevent train/test leakage, which THREE practices should they follow?

⚠ Common exam trap

The AI0-001 exam often tests the misconception that shuffling the entire dataset is always safe, but for temporal data or when duplicates exist, shuffling can introduce leakage by mixing future and past samples or spreading identical text across train and test sets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use time-based splitting for temporal data

Option B is correct because for temporal data, a time-based split (e.g., training on earlier timestamps and testing on later ones) prevents future information from leaking into the training set, which a random split would allow. Option C is correct because performing the train/test split before any cleaning, normalization, or other preprocessing ensures that statistics and transformations are learned only from the training data and not influenced by the test set. Option E is correct because removing duplicate samples and keeping all text from the same document in a single split prevents identical or near-identical content from appearing in both training and test sets, which would inflate performance estimates. Option A does not belong because shuffling the entire dataset before splitting is not inherently leakage-preventing and can actually cause leakage with temporal or grouped data. Option D does not belong because applying feature scaling to the entire dataset before splitting leaks test-set statistics (mean, variance) into training, which is a classic preprocessing leakage error.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Shuffle the entire dataset before splitting to ensure randomness

    Why it's wrong here

    Shuffling before splitting lets near-duplicate texts straddle the boundary, so test rows leak into training. Shuffling is correct after splitting, or when the dataset is already deduplicated and split by group. The stem's leakage requirement demands splitting first, then shuffling only the training partition.

  • ✓

    Use time-based splitting for temporal data

    Why this is correct

    Time-based splitting assigns earlier records to training and later records to testing, respecting chronological order. For temporal data this prevents future information leaking backwards into training, which random splitting would allow, thereby avoiding inflated evaluation results.

  • ✓

    Perform train/test split before any data cleaning or normalization

    Why this is correct

    Splitting before cleaning or normalisation ensures statistics such as means, vocabularies or scaling parameters are derived only from training data. Fitting these transformations on the full dataset would leak test-set information into training, biasing evaluation.

  • ✗

    Apply feature scaling to the entire dataset before splitting

    Why it's wrong here

    Fitting the scaler on the full dataset before splitting leaks test-set means and variances into training. Scaling belongs inside a pipeline fitted on training data only, then applied to the test partition. Whole-dataset scaling is right only when no held-out evaluation occurs, such as scoring production data.

  • ✓

    Remove duplicate samples and ensure that no text from the same document appears in both sets

    Why this is correct

    Duplicate samples and shared source documents create near-identical text in both splits, letting the model memorise test content during training. Deduplicating and grouping by document before splitting keeps the sets genuinely disjoint, preventing this leakage.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.