Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A data scientist is performing text preprocessing…

A data scientist is performing text preprocessing for a sentiment analysis model. The dataset contains many stop words and rare words. Which combination of preprocessing steps will reduce dimensionality and improve model performance?

⚠ Common exam trap

MLA-C01 often tests the misconception that any encoding scheme (label or one-hot) is a valid substitute for TF-IDF in text preprocessing, when in fact only TF-IDF captures term importance and reduces the impact of stop words.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Tokenization, stop-word removal, and TF-IDF

Tokenization splits raw text into individual terms, stop-word removal eliminates high-frequency but low-information words (e.g., 'the', 'is', 'and'), and TF-IDF assigns weights that downscale terms appearing in many documents while upscaling rare, discriminative terms. Together they reduce the feature space and emphasize sentiment-bearing vocabulary, which directly improves model performance on text classification tasks. This is the standard NLP preprocessing pipeline for sentiment analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove all words shorter than 3 characters and apply label encoding

    Why it's wrong here

    Label encoding assigns arbitrary integers to words, implying a false ordinal relationship and failing to reduce vocabulary dimensionality; removing stop words and applying TF-IDF or frequency thresholding achieves that. It is tempting because label encoding is standard for categorical targets, not for free-text tokens.

  • ✗

    Only tokenization without any removal

    Why it's wrong here

    Tokenisation alone retains every stop word and rare term as a distinct feature, leaving vocabulary size unchanged. It tempts because tokenisation is a mandatory first step, and would be correct only as part of a pipeline that also removes stop words and rare tokens.

  • ✓

    Tokenization, stop-word removal, and TF-IDF

    Why this is correct

    Tokenisation splits text into units, stop-word removal discards high-frequency, low-information terms, and TF-IDF down-weights terms appearing across many documents while rewarding rare, discriminative ones. Together these directly reduce the feature space created by the dataset's many stop words and rare words, satisfying the dimensionality-reduction constraint.

  • ✗

    Tokenization and one-hot encoding of words

    Why it's wrong here

    One-hot encoding creates a sparse column per unique token, so vocabulary size drives dimensionality upward rather than reducing it; rare words and stop words still each gain their own column. It is tempting because one-hot encoding suits small, fixed categorical vocabularies, where distinct labels need explicit binary representation.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.