Courseiva

AI0-001 AI Models and Data Engineering Practice Question

A data scientist is preparing a dataset for a natural language processing task. The dataset contains a 'review_text' column with free-form customer reviews. Before feeding the text into a machine learning model, the team wants to convert the text into numerical features. Which technique is most appropriate for this purpose?

⚠ Common exam trap

The trap here is thinking that any numerical conversion works, but techniques like hashing or average word length discard the semantic content needed for NLP tasks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Apply TF-IDF (Term Frequency-Inverse Document Frequency) vectorization to the 'review_text' column.

TF-IDF is the most appropriate because it transforms text into numerical vectors that reflect term importance, capturing semantic content while reducing the influence of common words. It is widely used for text classification and clustering, and it works well with many machine learning algorithms without requiring deep learning architectures.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use one-hot encoding on each unique word in the 'review_text' column.

    Why it's wrong here

    One-hot encoding each unique word would create a very high-dimensional and sparse matrix, one column per word, which is computationally expensive and prone to overfitting. It also ignores word importance and context. While it can be used for small vocabularies, it is not appropriate for free-form reviews with large vocabularies. TF-IDF is more efficient and informative.

  • ✗

    Convert each review to its average word length and use that as a single numerical feature.

    Why it's wrong here

    Average word length is a single scalar that loses almost all semantic information. It cannot capture the meaning or sentiment of the reviews, making it useless for most NLP tasks. This approach is far too simplistic and would result in poor model performance. It does not transform text into meaningful numerical features for machine learning.

  • ✓

    Apply TF-IDF (Term Frequency-Inverse Document Frequency) vectorization to the 'review_text' column.

    Why this is correct

    TF-IDF converts text into numerical vectors by weighting terms based on their frequency in a document and their rarity across the corpus. This highlights important words while downweighting common words like 'the' or 'and'. It is a standard and effective method for transforming unstructured text into features suitable for many machine learning algorithms, especially when the dataset is not extremely large.

  • ✗

    Hash each review to a fixed-length integer using a cryptographic hash function.

    Why it's wrong here

    Cryptographic hashing produces a fixed-length identifier that is not semantically meaningful and does not preserve any relationship between similar texts. It is used for security or deduplication, not for feature extraction. Using hashes as features would introduce noise and make it impossible for the model to learn patterns. It is not a valid text vectorization technique for NLP.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.