Courseiva
mediumMultiple Choice

Generative AI Leader Practice Question: Which Google AI model was the first to…

Which Google AI model was the first to demonstrate that transformers could be pre-trained bidirectionally on a large corpus, leading to major improvements in language understanding?

⚠ Common exam trap

In Google exams, candidates often confuse the original Transformer paper (introducing the architecture) with BERT's specific contribution of bidirectional pre-training, leading to selection of Option C instead of D.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

BERT

BERT (Bidirectional Encoder Representations from Transformers) was the first model to demonstrate that transformers could be pre-trained bidirectionally on a large corpus (BooksCorpus and English Wikipedia). By using a masked language model (MLM) objective, BERT conditions on both left and right context simultaneously, unlike previous unidirectional models, leading to significant improvements on 11 NLP benchmarks at its release.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    GPT-3

    Why it's wrong here

    GPT-3 is OpenAI's autoregressive, unidirectional decoder-only model trained for text generation, not bidirectional pre-training, and it is not a Google model. It is tempting as a well-known large language model, and would be correct for few-shot generation tasks, but not for the first bidirectional pre-training demonstration.

  • ✗

    AlphaGo

    Why it's wrong here

    AlphaGo is a reinforcement learning system that mastered the board game Go through self-play and Monte Carlo tree search; it performs no language pre-training. It is tempting as a landmark Google AI achievement, and would be correct for questions about game-playing agents or reinforcement learning milestones, not bidirectional language models.

  • ✗

    Transformer (the paper)

    Why it's wrong here

    The Transformer paper introduced the architecture and self-attention, but its encoder-decoder was trained for translation, not bidirectional masked pre-training on a large corpus. It is tempting because it is the foundational source of the architecture. BERT was the model that first demonstrated bidirectional pre-training for language understanding.

  • ✓

    BERT

    Why this is correct

    BERT pre-trains transformers bidirectionally using masked language modelling, so each token attends to both left and right context simultaneously. This bidirectional pre-training on a large corpus produced the language-understanding gains the stem describes, unlike unidirectional models such as GPT, which only attend to preceding tokens.

Visual reference

Client DHCP Server 1 Discover (broadcast) 2 Offer (IP: 192.168.1.10) 3 Request (I accept) 4 Acknowledge (lease confirmed) DORA — the four-step DHCP lease process

About these practice questions

One of 1,008 original Generative AI Leader practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This Generative AI Leader practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Generative AI Leader exam.