Courseiva
mediumMultiple Choice

MLA-C01 Practice Question: Building a sentiment analysis model for customer…

A company is building a sentiment analysis model for customer reviews. The text data contains many typos, abbreviations, and informal language. Which text preprocessing step would be most beneficial to support the ML model?

⚠ Common exam trap

MLA-C01 often tests the difference between generic preprocessing steps (stemming, stop-word removal) and the step that addresses the specific data quality problem — candidates may pick stemming because it is common, but it does not fix typos or abbreviations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply spell-check and normalize abbreviations

Typos, abbreviations, and informal language introduce noise that can cause the model to treat the same word as different tokens (e.g., 'gr8' vs 'great'), degrading sentiment classification. Applying spell-check and normalizing abbreviations standardizes the text so the model learns consistent patterns. This directly addresses the specific data quality issue described.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Apply stemming and lowercasing only

    Why it's wrong here

    Stemming and lowercasing normalise word forms but leave typos, abbreviations and slang untouched, so the model still sees inconsistent tokens. This suits clean, formal corpora where morphological reduction is the main need. Here, noisy informal text demands spelling correction and abbreviation expansion first.

  • ✗

    Remove all stop words

    Why it's wrong here

    Stop-word removal discards high-frequency words like 'not' and 'but', which carry sentiment polarity, and does nothing for typos or abbreviations. It suits topic classification or search indexing, where such tokens add noise rather than signal.

  • ✗

    Tokenize using whitespace tokenizer

    Why it's wrong here

    Whitespace tokenisation splits only on spaces, so it preserves typos, abbreviations and informal spellings as distinct tokens, leaving the noise the model must handle. It is tempting because it is fast and adequate for clean, well-formed text where vocabulary is already consistent. Here, normalisation or spelling correction is needed before tokenising.

  • ✓

    Apply spell-check and normalize abbreviations

    Why this is correct

    Spell-checking and abbreviation normalisation collapse noisy variants such as "gr8" and "thx" onto canonical tokens, shrinking vocabulary sparsity so the sentiment model learns consistent signal from informal reviews rather than treating each typo as a distinct feature.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.