Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: A team is preparing text data for sentiment…

A team is preparing text data for sentiment analysis. They have a large corpus of customer reviews. They want to convert the text into numerical features using a technique that captures word importance relative to the whole corpus. Which feature extraction method should they use?

⚠ Common exam trap

MLA-C01 often tests the distinction between frequency-based methods (CountVectorizer, TF-IDF) and embedding-based methods (Word2Vec) — candidates pick Word2Vec for 'importance' when the question specifically asks for corpus-relative term weighting.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

TF-IDF

TF-IDF (Term Frequency–Inverse Document Frequency) weights each word by how often it appears in a document relative to how often it appears across the entire corpus, so common words like 'the' are down-weighted while distinctive words are up-weighted. This directly captures word importance relative to the whole corpus, which is exactly what the team needs for sentiment analysis feature extraction.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Word2Vec embeddings

    Why it's wrong here

    Word2Vec produces dense semantic embeddings from local context windows, not corpus-relative importance weights, so it does not satisfy the requirement to capture word importance across the whole corpus. It is tempting because embeddings capture meaning, but TF-IDF is the correct choice when term weighting by document frequency is required.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding assigns each token a binary column with no weighting, so it cannot express a word's importance relative to the corpus and produces sparse, high-dimensional vectors. It is tempting for small vocabularies or categorical labels, but TF-IDF is the method that scales term frequency by inverse document frequency across the corpus.

  • ✓

    TF-IDF

    Why this is correct

    TF-IDF weights each term by its frequency within a review against its rarity across the whole corpus, so words that distinguish individual reviews score higher than common ones. That corpus-relative importance weighting is exactly what the scenario requires.

  • ✗

    Bag-of-words (CountVectorizer)

    Why it's wrong here

    CountVectorizer records raw term frequencies without inverse document frequency, so common words across the corpus are not down-weighted and importance relative to the whole corpus is not captured. It is tempting as a simple baseline, but TF-IDF applies the IDF weighting the scenario explicitly requires.

About these practice questions

One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.