Courseiva
easyMultiple Choice

AIF-C01 Practice Question: The primary advantage of the transformer…

What is the primary advantage of the transformer architecture over previous RNN-based architectures for natural language processing tasks?

⚠ Common exam trap

In the AWS AI Practitioner exam, a common trap is to confuse the sequential processing of RNNs with a desirable property for maintaining order, but transformers actually use positional encodings to preserve sequence order while still enabling parallel processing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It relies on a self-attention mechanism that allows parallel processing of all tokens

The transformer architecture's primary advantage is its self-attention mechanism, which allows the model to process all tokens in the input sequence in parallel rather than sequentially. This parallelization dramatically reduces training time and enables the model to capture long-range dependencies more effectively than RNNs, which must process tokens one by one and suffer from vanishing gradients over long sequences.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It processes tokens sequentially to maintain the order of the sequence

    Why it's wrong here

    Transformers process all tokens in parallel through self-attention, capturing long-range dependencies without sequential recurrence. Sequential processing is the RNN behaviour transformers replaced; retaining it would reintroduce the vanishing-gradient and slow-training limits that attention was designed to remove.

  • ✗

    It requires less training data because it uses convolutional layers

    Why it's wrong here

    Self-attention lets transformers model long-range dependencies in parallel, and they typically require more data than RNNs, not less. Convolutional layers belong to CNNs, which capture local spatial patterns; transformers use attention instead, so this describes neither architecture's advantage.

  • ✗

    It uses a gated recurrent unit to selectively forget information

    Why it's wrong here

    GRUs are themselves recurrent, so this describes an older RNN variant rather than the transformer's advantage. It is tempting because gating does address vanishing gradients in sequence models, but the transformer's actual gain is self-attention, which processes all tokens in parallel instead of sequentially.

  • ✓

    It relies on a self-attention mechanism that allows parallel processing of all tokens

    Why this is correct

    Self-attention computes relationships between every token pair simultaneously, so the model processes all positions in parallel rather than sequentially as RNNs must. This removes the sequential dependency that prevented parallelisation, directly satisfying the stem's demand for the primary advantage: dramatically faster training on long sequences.

About these practice questions

Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.