NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A research team is pretraining a transformer on a corpus of 200 billion tokens. They want the model to learn bidirectional context so each token attends to both left and right neighbors during pretraining. Which pretraining objective fits this requirement?
⚠ Common exam trap
The trap here is equating any language modeling objective with bidirectional context, when causal masking in autoregressive objectives blocks attention to future tokens.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Masked language modeling
Bidirectional context means each token's representation is informed by both preceding and following tokens. Masked language modeling achieves this by hiding random tokens and requiring the model to reconstruct them from the full unmasked context, so attention spans the entire sequence. Autoregressive and causal-decoder objectives enforce left-to-right masking, which precludes bidirectional attention.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Masked language modeling
Why this is correct
Masked language modeling randomly replaces a fraction of input tokens with a mask and trains the model to recover them using the full surrounding context on both sides. Because no causal mask is applied, every prediction can attend to tokens to the left and right, which is exactly the bidirectional pretraining signal the team is asking for.
- ✗
Contrastive next-sentence prediction
Why it's wrong here
Contrastive next-sentence prediction trains the model to distinguish whether one sentence follows another, producing sentence-level representations rather than token-level bidirectional understanding. It was a secondary objective in early bidirectional encoders and does not by itself teach each token to attend to both neighbors across the corpus.
- ✗
Sequence-to-sequence denoising with a causal decoder
Why it's wrong here
A causal decoder in a sequence-to-sequence setup still applies a left-to-right attention mask during generation, so the decoder cannot see future tokens. While the encoder may be bidirectional, the team's stated goal of bidirectional context for every token during pretraining is not achieved by this objective alone.
- ✗
Autoregressive next-token prediction
Why it's wrong here
Autoregressive next-token prediction uses a causal mask that prevents each position from attending to future tokens, so the model only sees left context. This is the objective used by decoder-only models such as GPT, and it cannot provide the bidirectional context the team requires during pretraining.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.