mediumMultiple Choice
MLA-C01 Practice Question: A data scientist is performing text preprocessing…
A data scientist is performing text preprocessing for a sentiment analysis model. The dataset contains many stop words and rare words. Which combination of preprocessing steps will reduce dimensionality and improve model performance?
⚠ Common exam trap
MLA-C01 often tests the misconception that any encoding scheme (label or one-hot) is a valid substitute for TF-IDF in text preprocessing, when in fact only TF-IDF captures term importance and reduces the impact of stop words.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tokenization, stop-word removal, and TF-IDF
Tokenization splits raw text into individual terms, stop-word removal eliminates high-frequency but low-information words (e.g., 'the', 'is', 'and'), and TF-IDF assigns weights that downscale terms appearing in many documents while upscaling rare, discriminative terms. Together they reduce the feature space and emphasize sentiment-bearing vocabulary, which directly improves model performance on text classification tasks. This is the standard NLP preprocessing pipeline for sentiment analysis.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove all words shorter than 3 characters and apply label encoding
Why it's wrong here
Label encoding assigns arbitrary integers to words, implying a false ordinal relationship and failing to reduce vocabulary dimensionality; removing stop words and applying TF-IDF or frequency thresholding achieves that. It is tempting because label encoding is standard for categorical targets, not for free-text tokens.
- ✗
Only tokenization without any removal
Why it's wrong here
Tokenisation alone retains every stop word and rare term as a distinct feature, leaving vocabulary size unchanged. It tempts because tokenisation is a mandatory first step, and would be correct only as part of a pipeline that also removes stop words and rare tokens.
- ✓
Tokenization, stop-word removal, and TF-IDF
Why this is correct
Tokenisation splits text into units, stop-word removal discards high-frequency, low-information terms, and TF-IDF down-weights terms appearing across many documents while rewarding rare, discriminative ones. Together these directly reduce the feature space created by the dataset's many stop words and rare words, satisfying the dimensionality-reduction constraint.
- ✗
Tokenization and one-hot encoding of words
Why it's wrong here
One-hot encoding creates a sparse column per unique token, so vocabulary size drives dimensionality upward rather than reducing it; rare words and stop words still each gain their own column. It is tempting because one-hot encoding suits small, fixed categorical vocabularies, where distinct labels need explicit binary representation.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.