NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A research team is designing a decoder-only transformer LLM for long-document question answering. They want to reduce the quadratic computational cost of self-attention so that training on sequences of 32,000 tokens is feasible on their GPU cluster. Which two techniques are appropriate for this goal? (Choose two.)
⚠ Common exam trap
A common mix-up: candidates confuse general model-size or architecture changes, such as more heads or a larger vocabulary, with techniques that actually reduce the quadratic scaling of attention over sequence length.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Sparse attention patterns that restrict each token to a subset of positions
Sparse attention lowers the number of attention interactions, and IO-aware exact attention kernels like FlashAttention reduce memory traffic and speed up the exact computation. Both directly target the quadratic cost of self-attention for long sequences. Increasing heads, enlarging the vocabulary, or swapping normalization layers does not change the sequence-length-squared scaling and therefore does not make 32,000-token training feasible.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Using a larger vocabulary size for the tokenizer
Why it's wrong here
Vocabulary size affects embedding parameters and how many tokens represent a given text, but it does not alter the quadratic cost of self-attention. In fact, a larger vocabulary can increase parameters and memory. While it may slightly shorten sequences by merging common words, it is not a technique for reducing attention complexity and is not appropriate for the stated goal.
- ✓
Sparse attention patterns that restrict each token to a subset of positions
Why this is correct
Sparse attention reduces the number of query-key interactions from quadratic to near-linear by having each token attend to a limited set of positions, such as a local window plus selected global tokens. This directly lowers compute and memory for long sequences. It is a standard approach for making 32,000-token training feasible while retaining much of the model's ability to capture long-range dependencies.
- ✗
Replacing layer normalization with batch normalization in the transformer blocks
Why it's wrong here
Batch normalization is poorly suited to variable-length sequence models and does not reduce the quadratic attention cost. Layer normalization is standard in transformers because it normalizes across features per token, independent of batch statistics. Swapping it would harm training stability and do nothing to make 32,000-token sequences computationally feasible.
- ✓
FlashAttention-style IO-aware exact attention kernels
Why this is correct
FlashAttention computes exact attention but reorganizes the computation to minimize reads and writes to high-bandwidth memory, which substantially speeds up training and reduces memory usage for long sequences. It does not approximate the attention matrix, so quality is preserved. This makes it a practical technique for training on 32,000-token sequences without changing the model architecture.
- ✗
Increasing the number of attention heads while keeping the same total hidden dimension
Why it's wrong here
Adding more attention heads with a fixed hidden dimension reduces the per-head dimension but does not change the quadratic scaling of attention with sequence length. Total attention compute remains proportional to sequence length squared. This architectural tweak may affect representational capacity, but it does not address the core computational bottleneck for 32,000-token sequences.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.