Courseiva
mediumMultiple Select

MLA-C01 Practice Question: A data scientist is preparing text data for a…

A data scientist is preparing text data for a sentiment analysis model using Amazon SageMaker. Which two data preprocessing techniques are commonly used when working with text data for natural language processing? (Choose two.)

⚠ Common exam trap

Watch out — candidates often confuse one-hot encoding as a preprocessing technique for raw text, when it is actually a feature engineering step applied after tokenization, and they may overlook that stop word removal is a standard preprocessing step despite its potential to remove sentiment-bearing words in certain contexts.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Tokenization

Tokenization (C) is correct because NLP models require raw text to be split into individual tokens (words, subwords, or characters) so they can be mapped to numerical vectors before being fed into a sentiment analysis model in SageMaker. Stop word removal (E) is correct because eliminating high-frequency, low-information words such as 'the', 'is', and 'and' reduces noise and dimensionality, helping the model focus on sentiment-bearing terms. One-hot encoding of all words (A) is not a common standalone preprocessing technique for text at scale, since it produces extremely sparse, high-dimensional vectors and ignores word order and semantics; embeddings are typically preferred. Image resizing (B) applies to computer vision data, not text. Principal component analysis (D) is a dimensionality reduction technique for numerical feature matrices, not a standard text preprocessing step for NLP.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding of all words

    Why it's wrong here

    One-hot encoding produces a sparse, high-dimensional vector per word and discards word order and context, so it cannot capture the sequential dependencies sentiment analysis needs. It is tempting because it is a genuine encoding technique, and it would suit categorical features with a small, fixed vocabulary rather than free text.

  • ✗

    Image resizing

    Why it's wrong here

    Image resizing alters pixel dimensions of visual data, so it cannot tokenise, normalise or vectorise text for a sentiment model. It is tempting because it is a standard preprocessing step, and it would be the correct choice when preparing image data for a computer vision model instead.

  • ✓

    Tokenization

    Why this is correct

    Tokenization splits raw text into individual words or subword units, converting unstructured strings into discrete tokens that models can vectorise. This satisfies the text-preprocessing requirement, forming the essential first step before embedding or vectorisation in NLP pipelines.

  • ✗

    Principal component analysis (PCA)

    Why it's wrong here

    PCA is a dimensionality reduction technique for numerical data, not typical for text preprocessing.

  • ✓

    Stop word removal

    Why this is correct

    Stop word removal strips high-frequency, low-information tokens such as "the" and "and" from the corpus, reducing vocabulary size and noise so the sentiment model focuses on semantically meaningful terms. This directly satisfies the question's requirement for a common NLP text preprocessing technique alongside tokenisation and stemming.

About these practice questions

This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.