easyMultiple Choice
Generative AI Leader Practice Question: The primary purpose of the transformer…
What is the primary purpose of the transformer architecture in large language models (LLMs)?
⚠ Common exam trap
Test-takers frequently confuse the transformer's core innovation (parallel self-attention for sequence modeling) with auxiliary tasks like embedding generation or retrieval, which are separate components in an LLM pipeline, leading them to pick C or D as plausible but incorrect answers.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
To enable parallel processing of tokens and capture long-range dependencies through self-attention
The transformer architecture's primary purpose is to enable parallel processing of all tokens in a sequence while capturing long-range dependencies through its self-attention mechanism. Unlike recurrent neural networks (RNNs) that process tokens sequentially, transformers compute attention scores between every pair of tokens simultaneously, allowing the model to weigh the relevance of distant tokens without the vanishing gradient problem. This parallelization and global context capture are the foundational innovations that make large language models (LLMs) scalable and effective for tasks like text generation and understanding.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
To generate images from text descriptions
Why it's wrong here
Transformers process token sequences via self-attention to model language; image generation is performed by diffusion or multimodal decoders, not the architecture itself. It is tempting because transformer blocks do underpin many text-to-image systems, so the architecture appears in that pipeline even though generation is not its purpose.
- ✓
To enable parallel processing of tokens and capture long-range dependencies through self-attention
Why this is correct
Self-attention lets every token attend to every other token simultaneously, so the model captures long-range dependencies without sequential recurrence. This parallel processing across the whole sequence is what makes transformers trainable at the scale LLMs require, unlike RNN architectures that process tokens one at a time.
- ✗
To convert text into numerical embeddings for downstream tasks
Why it's wrong here
Embeddings come from separate encoder models or embedding layers, not the transformer's defining function; the architecture's purpose is modelling token relationships through self-attention to predict or generate text. It is tempting because transformers do contain embedding layers, and encoder-only variants such as BERT produce embeddings for downstream tasks.
- ✗
To store and retrieve information from a vector database
Why it's wrong here
Vector databases are external retrieval stores; transformers compute attention over input tokens and hold no persistent storage or retrieval mechanism. It is tempting because retrieval-augmented generation pairs a vector database with a transformer, so the two appear together, but storage is the database's role, not the architecture's.
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.