mediumMultiple Select
MLA-C01 Practice Question: A data scientist is preparing text data for a…
A data scientist is preparing text data for a sentiment analysis model using Amazon SageMaker. Which two data preprocessing techniques are commonly used when working with text data for natural language processing? (Choose two.)
⚠ Common exam trap
Watch out — candidates often confuse one-hot encoding as a preprocessing technique for raw text, when it is actually a feature engineering step applied after tokenization, and they may overlook that stop word removal is a standard preprocessing step despite its potential to remove sentiment-bearing words in certain contexts.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Tokenization
Tokenization (C) is correct because NLP models require raw text to be split into individual tokens (words, subwords, or characters) so they can be mapped to numerical vectors before being fed into a sentiment analysis model in SageMaker. Stop word removal (E) is correct because eliminating high-frequency, low-information words such as 'the', 'is', and 'and' reduces noise and dimensionality, helping the model focus on sentiment-bearing terms. One-hot encoding of all words (A) is not a common standalone preprocessing technique for text at scale, since it produces extremely sparse, high-dimensional vectors and ignores word order and semantics; embeddings are typically preferred. Image resizing (B) applies to computer vision data, not text. Principal component analysis (D) is a dimensionality reduction technique for numerical feature matrices, not a standard text preprocessing step for NLP.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encoding of all words
Why it's wrong here
One-hot encoding produces a sparse, high-dimensional vector per word and discards word order and context, so it cannot capture the sequential dependencies sentiment analysis needs. It is tempting because it is a genuine encoding technique, and it would suit categorical features with a small, fixed vocabulary rather than free text.
- ✗
Image resizing
Why it's wrong here
Image resizing alters pixel dimensions of visual data, so it cannot tokenise, normalise or vectorise text for a sentiment model. It is tempting because it is a standard preprocessing step, and it would be the correct choice when preparing image data for a computer vision model instead.
- ✓
Tokenization
Why this is correct
Tokenization splits raw text into individual words or subword units, converting unstructured strings into discrete tokens that models can vectorise. This satisfies the text-preprocessing requirement, forming the essential first step before embedding or vectorisation in NLP pipelines.
- ✗
Principal component analysis (PCA)
Why it's wrong here
PCA is a dimensionality reduction technique for numerical data, not typical for text preprocessing.
- ✓
Stop word removal
Why this is correct
Stop word removal strips high-frequency, low-information tokens such as "the" and "and" from the corpus, reducing vocabulary size and noise so the sentiment model focuses on semantically meaningful terms. This directly satisfies the question's requirement for a common NLP text preprocessing technique alongside tokenisation and stemming.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.