Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A data scientist is preparing text data for…

A data scientist is preparing text data for sentiment analysis. They need to convert the text into numerical features while reducing the impact of common words. Which feature extraction method should they use?

⚠ Common exam trap

MLA-C01 often tests the confusion between CountVectorizer (raw counts) and TF-IDF (weighted counts) — candidates pick D, but only TF-IDF reduces the impact of common words via inverse document frequency.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

TF-IDF vectorization

TF-IDF (Term Frequency–Inverse Document Frequency) converts text to numerical features while down-weighting common words (like 'the', 'is') that appear across many documents, exactly matching the requirement to reduce the impact of common words. It is the standard feature extraction method for sentiment analysis when common-word influence must be minimized.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Word2Vec embeddings

    Why it's wrong here

    Word2Vec produces dense contextual embeddings that capture semantic similarity but assign no inverse-document-frequency weighting, so common words are not down-weighted. It is tempting because it is the standard neural text representation, and would be correct when semantic relationships between words matter more than reducing frequent-term influence.

  • ✓

    TF-IDF vectorization

    Why this is correct

    TF-IDF vectorization weights each term by its frequency within a document, scaled inversely by how many documents contain it. This downweights common words such as "the" while emphasising distinctive terms, directly satisfying the requirement to reduce the impact of frequent words during numerical feature extraction.

  • ✗

    Label encoding of each word

    Why it's wrong here

    Label encoding assigns each distinct word an arbitrary integer, producing ordinal values with no frequency weighting, so common words retain full influence and no document-term matrix is built. It is tempting because it converts text to numbers cheaply, and would suit encoding a fixed set of categorical labels rather than free-text sentiment features.

  • ✗

    CountVectorizer with n-grams

    Why it's wrong here

    CountVectorizer with n-grams counts raw term occurrences, so frequent common words dominate the feature matrix unless stop words or TF-IDF weighting are applied separately. It is tempting because it is the standard bag-of-words extraction method, and would be correct when term frequency alone is wanted without down-weighting common words.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.