mediumMultiple Choice
Generative AI Leader Practice Question: Which Google AI model was the first to…
Which Google AI model was the first to demonstrate that transformers could be pre-trained bidirectionally on a large corpus, leading to major improvements in language understanding?
⚠ Common exam trap
In Google exams, candidates often confuse the original Transformer paper (introducing the architecture) with BERT's specific contribution of bidirectional pre-training, leading to selection of Option C instead of D.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
BERT
BERT (Bidirectional Encoder Representations from Transformers) was the first model to demonstrate that transformers could be pre-trained bidirectionally on a large corpus (BooksCorpus and English Wikipedia). By using a masked language model (MLM) objective, BERT conditions on both left and right context simultaneously, unlike previous unidirectional models, leading to significant improvements on 11 NLP benchmarks at its release.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
GPT-3
Why it's wrong here
GPT-3 is OpenAI's autoregressive, unidirectional decoder-only model trained for text generation, not bidirectional pre-training, and it is not a Google model. It is tempting as a well-known large language model, and would be correct for few-shot generation tasks, but not for the first bidirectional pre-training demonstration.
- ✗
AlphaGo
Why it's wrong here
AlphaGo is a reinforcement learning system that mastered the board game Go through self-play and Monte Carlo tree search; it performs no language pre-training. It is tempting as a landmark Google AI achievement, and would be correct for questions about game-playing agents or reinforcement learning milestones, not bidirectional language models.
- ✗
Transformer (the paper)
Why it's wrong here
The Transformer paper introduced the architecture and self-attention, but its encoder-decoder was trained for translation, not bidirectional masked pre-training on a large corpus. It is tempting because it is the foundational source of the architecture. BERT was the model that first demonstrated bidirectional pre-training for language understanding.
- ✓
BERT
Why this is correct
BERT pre-trains transformers bidirectionally using masked language modelling, so each token attends to both left and right context simultaneously. This bidirectional pre-training on a large corpus produced the language-understanding gains the stem describes, unlike unidirectional models such as GPT, which only attend to preceding tokens.
Visual reference
Go deeper
Related to this question
About these practice questions
One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.