Courseiva

AIF-C01 Fundamentals of Generative AI Practice Question

A data science intern is reading about the transformer architecture that underpins most modern large language models. The intern asks which mechanism allows a transformer to weigh the relevance of every other word in a sentence when encoding a given word, regardless of how far apart the words appear. Which mechanism should you name?

⚠ Common exam trap

The trap here is assuming that positional encoding is what gives a transformer its long-range context, when positional encoding only supplies order information and self-attention does the actual cross-token weighting.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Self-attention

Self-attention is the operation that lets each token build its representation from all other tokens in the sequence with learned weights, which is what removes the long-distance constraint of sequential models. Positional encoding, tokenization, and recurrent state propagation all play supporting or unrelated roles, but none of them performs the relevance weighting described in the scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Self-attention

    Why this is correct

    Self-attention computes a weighted relationship between every token and every other token in the sequence, so a word's representation is built from the whole context rather than a fixed local window. This directly removes the distance limitation of recurrent architectures, which is exactly the property the intern is asking about, and it is the core building block of the transformer encoder and decoder stacks.

  • ✗

    Positional encoding

    Why it's wrong here

    Positional encoding injects order information into token embeddings because attention itself is order-agnostic. It is a necessary companion to attention, but it does not compute relevance weights between words. Naming it confuses the mechanism that supplies sequence position with the mechanism that mixes information across positions, which is the distinction the scenario is probing.

  • ✗

    Tokenization

    Why it's wrong here

    Tokenization converts raw text into the subword units a model consumes, and it happens before any representation is computed. It determines vocabulary and sequence length but performs no weighting of one token against another, so it cannot explain how distant words influence each other inside a transformer layer.

  • ✗

    Recurrent state propagation

    Why it's wrong here

    Recurrent state propagation carries information step by step through a hidden state, which forces sequential processing and weakens the influence of distant tokens over long sequences. It describes RNN and LSTM style models, not the parallelism and global context mixing that transformers are known for, so it does not answer what lets a transformer attend to any position in the sentence.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.