Courseiva
mediumMultiple Choice

AIF-C01 Practice Question: A data scientist wants to compare the text…

A data scientist wants to compare the text embeddings generated by Amazon Titan Embeddings for a set of product descriptions. Which metric is MOST appropriate to measure the semantic similarity between two embedding vectors?

⚠ Common exam trap

A common mistake is to assume Euclidean distance works for text embeddings, but AWS Titan Embeddings produces unit vectors, so cosine similarity is the standard metric.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cosine similarity

Cosine similarity measures the cosine of the angle between two vectors, focusing on their orientation rather than magnitude. For text embeddings from models like Amazon Titan Embeddings, which are normalized to unit length, cosine similarity is the standard metric because it captures semantic similarity even when descriptions differ in length or word count. Euclidean distance would be affected by vector magnitude, making it less suitable for comparing semantic content.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Hamming distance

    Why it's wrong here

    Hamming distance counts differing positions in equal-length strings, which is meaningless for dense floating-point embeddings. It is tempting because it is a genuine distance metric, and it would be correct for comparing fixed-length binary codes or error-corrected bitstrings.

  • ✗

    Manhattan distance

    Why it's wrong here

    Manhattan distance sums absolute coordinate differences and ignores vector direction, so it poorly reflects semantic similarity. It is tempting because it is a valid distance metric, and it would be correct for grid-like or sparse feature spaces where axis-aligned differences matter.

  • ✗

    Euclidean distance

    Why it's wrong here

    Euclidean distance measures raw geometric separation, so it is sensitive to vector magnitude and fails to isolate the directional similarity that cosine similarity captures for Titan embeddings. It is tempting because it is the standard distance metric in low-dimensional clustering, where absolute position genuinely matters.

  • ✓

    Cosine similarity

    Why this is correct

    Cosine similarity measures the angle between two vectors, ignoring magnitude, so it reflects semantic orientation rather than length. This suits comparing Titan Embeddings, where semantic closeness is encoded in direction, and it remains stable across vectors of differing dimensionality-normalised scales.

About these practice questions

This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.