Courseiva

AI0-001 AI Concepts and Foundations Practice Question

A team is building a natural language processing (NLP) model to analyze customer feedback. They have a large corpus of unlabeled text data and want to generate word embeddings that capture semantic meaning. Which approach should they use?

⚠ Common exam trap

CompTIA often tests the distinction between frequency-based vectorization (TF-IDF, bag-of-words) and prediction-based embedding methods (Word2Vec, GloVe), trapping candidates who think TF-IDF captures semantic meaning when it only captures term importance in a document.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Word2Vec

Word2Vec is the correct approach because it learns dense, distributed word embeddings from large unlabeled corpora by training a shallow neural network to predict words in context (CBOW) or context from words (Skip-gram). This captures semantic relationships such as analogy and similarity, which is essential for analyzing customer feedback without labeled data.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding produces sparse orthogonal vectors with no shared structure, so it cannot capture semantic similarity between words. It is tempting as a basic text-encoding step, and it would be correct for representing categorical variables or small vocabularies in simple models.

  • ✗

    TF-IDF vectorization

    Why it's wrong here

    TF-IDF weights terms by frequency and rarity, yielding sparse vectors whose dimensions remain independent, so semantic relationships are not encoded. It is tempting because it is a standard text-vectorisation technique, and it would be correct for document classification or keyword retrieval tasks.

  • ✓

    Word2Vec

    Why this is correct

    Word2Vec learns dense vector representations from unlabelled text by predicting a word from its neighbours (skip-gram) or vice versa (CBOW), so semantic relationships emerge from co-occurrence statistics. This directly satisfies the stem's requirement for embeddings from a large unlabelled corpus, unlike supervised approaches needing labelled data.

  • ✗

    Bag-of-words model

    Why it's wrong here

    Bag-of-words counts token occurrences and discards order and context, producing sparse vectors with no semantic relationships between terms. It is tempting as a simple text representation, and it would be correct for basic sentiment classification or topic detection with small datasets.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.