AI0-001 Machine Learning and Deep Learning Practice Question
A data science team is preparing a dataset of customer support tickets to train a supervised model that routes each ticket to the correct department. They have 40,000 tickets labeled with one of eight departments. Which TWO preprocessing steps are most appropriate before training? (Choose two.)
⚠ Common exam trap
The trap here is fitting the vectorizer on the entire dataset before splitting, which leaks vocabulary and statistics from the test set into training and inflates reported accuracy.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Split the data into training, validation, and test sets before fitting any preprocessing.
Text classification requires converting raw strings into numeric features and preserving an honest evaluation split. Tokenization with TF-IDF or embeddings supplies the numerical representation, while splitting before fitting preprocessing prevents leakage. Together these steps prepare the ticket data so the routing model can be trained and fairly assessed on unseen examples.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Split the data into training, validation, and test sets before fitting any preprocessing.
Why this is correct
Holding out validation and test data before fitting vectorizers or scalers prevents information leakage from the evaluation sets into training. If TF-IDF vocabulary or normalization statistics are learned on the full dataset, reported performance becomes optimistic and unreliable. This step ensures the routing model is evaluated on genuinely unseen tickets.
- ✗
Delete all tickets shorter than 20 words to reduce noise.
Why it's wrong here
Short tickets often contain clear routing signals such as 'password reset' or 'billing error', so deleting them removes valid training examples and can bias the model. Length is not a reliable indicator of label quality. A better approach is to review a sample of short tickets and clean only those that are genuinely malformed.
- ✗
Standardize the department labels to zero mean and unit variance.
Why it's wrong here
Department labels are categorical classes, not continuous values, so computing means and variances on them is meaningless. Standardization applies to numerical features, not class labels. The correct treatment for the eight department labels is encoding them as class indices or one-hot targets, not statistical normalization.
- ✓
Tokenize the ticket text and convert it to numerical vectors using an embedding or TF-IDF representation.
Why this is correct
Machine learning models require numerical input, and ticket text is unstructured. Tokenization splits text into units, and TF-IDF or embeddings convert those units into vectors that preserve semantic or statistical information. This is a required step for any text classification pipeline and directly supports routing tickets to the correct department.
- ✗
Apply one-hot encoding to the raw ticket text strings.
Why it's wrong here
One-hot encoding creates a column per unique token, which for 40,000 tickets yields a huge, extremely sparse matrix with little semantic value. It also fails to capture similarity between related words like 'refund' and 'reimbursement'. Vectorization methods such as TF-IDF or embeddings are designed for text and are far more effective here.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.