Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

A research team is pretraining a transformer on a corpus of 200 billion tokens. They want the model to learn bidirectional context so each token attends to both left and right neighbors during pretraining. Which pretraining objective fits this requirement?

⚠ Common exam trap

The trap here is equating any language modeling objective with bidirectional context, when causal masking in autoregressive objectives blocks attention to future tokens.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Masked language modeling

Bidirectional context means each token's representation is informed by both preceding and following tokens. Masked language modeling achieves this by hiding random tokens and requiring the model to reconstruct them from the full unmasked context, so attention spans the entire sequence. Autoregressive and causal-decoder objectives enforce left-to-right masking, which precludes bidirectional attention.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Masked language modeling

    Why this is correct

    Masked language modeling randomly replaces a fraction of input tokens with a mask and trains the model to recover them using the full surrounding context on both sides. Because no causal mask is applied, every prediction can attend to tokens to the left and right, which is exactly the bidirectional pretraining signal the team is asking for.

  • ✗

    Contrastive next-sentence prediction

    Why it's wrong here

    Contrastive next-sentence prediction trains the model to distinguish whether one sentence follows another, producing sentence-level representations rather than token-level bidirectional understanding. It was a secondary objective in early bidirectional encoders and does not by itself teach each token to attend to both neighbors across the corpus.

  • ✗

    Sequence-to-sequence denoising with a causal decoder

    Why it's wrong here

    A causal decoder in a sequence-to-sequence setup still applies a left-to-right attention mask during generation, so the decoder cannot see future tokens. While the encoder may be bidirectional, the team's stated goal of bidirectional context for every token during pretraining is not achieved by this objective alone.

  • ✗

    Autoregressive next-token prediction

    Why it's wrong here

    Autoregressive next-token prediction uses a causal mask that prevents each position from attending to future tokens, so the model only sees left context. This is the objective used by decoder-only models such as GPT, and it cannot provide the bidirectional context the team requires during pretraining.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.