easyMultiple Choice
AIF-C01 Practice Question: Which component of the Transformer architecture…
Which component of the Transformer architecture allows the model to weigh the importance of different tokens in the input sequence when generating each output token?
⚠ Common exam trap
AWS often tests the misconception that positional encoding or feed-forward networks handle token relationships, but the self-attention mechanism is the only component that directly computes pairwise token importance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Self-attention mechanism
The self-attention mechanism (option D) is the core component of the Transformer that computes attention scores between every pair of tokens in the input sequence, allowing the model to dynamically weigh the importance of each token when generating an output token. This enables the model to capture long-range dependencies and contextual relationships without the sequential constraints of RNNs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Layer normalization
Why it's wrong here
Layer normalization stabilises activations between sublayers, keeping training numerically stable; it computes no token-to-token weighting. Self-attention produces those attention weights. Layer normalization is tempting because it appears in every Transformer block, and it would be correct when addressing unstable gradients or slow convergence during training.
- ✗
Feed-forward network
Why it's wrong here
The feed-forward network transforms each token's representation independently after attention has mixed information across positions; it performs no weighting between tokens. Self-attention does that weighting. The feed-forward network is tempting because it holds most parameters, and it would be correct when describing where per-token non-linear transformation occurs.
- ✗
Positional encoding
Why it's wrong here
Positional encoding injects sequence-order information into token embeddings so the model knows where each token sits; it assigns no importance weights. Self-attention computes those weights. Positional encoding is tempting because Transformers lack recurrence, and it would be correct when explaining how word order is preserved despite parallel processing.
- ✓
Self-attention mechanism
Why this is correct
Self-attention computes query, key and value projections for every token, then scores each token against all others so their weighted contributions shape the current output. This weighting is the mechanism the stem requires for judging token importance, unlike feed-forward layers, which process positions independently.
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.