mediumMultiple Choice
AIF-C01 Practice Question: A data scientist wants to compare the text…
A data scientist wants to compare the text embeddings generated by Amazon Titan Embeddings for a set of product descriptions. Which metric is MOST appropriate to measure the semantic similarity between two embedding vectors?
⚠ Common exam trap
A common mistake is to assume Euclidean distance works for text embeddings, but AWS Titan Embeddings produces unit vectors, so cosine similarity is the standard metric.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cosine similarity
Cosine similarity measures the cosine of the angle between two vectors, focusing on their orientation rather than magnitude. For text embeddings from models like Amazon Titan Embeddings, which are normalized to unit length, cosine similarity is the standard metric because it captures semantic similarity even when descriptions differ in length or word count. Euclidean distance would be affected by vector magnitude, making it less suitable for comparing semantic content.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Hamming distance
Why it's wrong here
Hamming distance counts differing positions in equal-length strings, which is meaningless for dense floating-point embeddings. It is tempting because it is a genuine distance metric, and it would be correct for comparing fixed-length binary codes or error-corrected bitstrings.
- ✗
Manhattan distance
Why it's wrong here
Manhattan distance sums absolute coordinate differences and ignores vector direction, so it poorly reflects semantic similarity. It is tempting because it is a valid distance metric, and it would be correct for grid-like or sparse feature spaces where axis-aligned differences matter.
- ✗
Euclidean distance
Why it's wrong here
Euclidean distance measures raw geometric separation, so it is sensitive to vector magnitude and fails to isolate the directional similarity that cosine similarity captures for Titan embeddings. It is tempting because it is the standard distance metric in low-dimensional clustering, where absolute position genuinely matters.
- ✓
Cosine similarity
Why this is correct
Cosine similarity measures the angle between two vectors, ignoring magnitude, so it reflects semantic orientation rather than length. This suits comparing Titan Embeddings, where semantic closeness is encoded in direction, and it remains stable across vectors of differing dimensionality-normalised scales.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.