NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A research team is training a transformer-based language model and wants to reduce the computational cost of the self-attention mechanism for very long input sequences. They are considering replacing the standard scaled dot-product attention with an approximation. Which statement accurately describes a trade-off of using an approximate attention method?
⚠ Common exam trap
The trap here is assuming approximate attention only adds benefits without drawbacks, when in fact it trades some model quality for reduced memory and compute on long sequences.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It reduces memory and compute complexity but may sacrifice some model quality compared to full attention.
Standard self-attention scales quadratically with sequence length, which is costly for long inputs. Approximate attention methods reduce this complexity to near-linear or sub-quadratic by sparsifying or factorizing the attention matrix. The trade-off is that the approximation may not capture all token interactions, potentially lowering model quality or requiring additional training. This balance between efficiency and accuracy is the central consideration.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
It reduces memory and compute complexity but may sacrifice some model quality compared to full attention.
Why this is correct
Approximate attention methods, such as sparse or low-rank attention, lower the quadratic complexity of full attention to near-linear or sub-quadratic, saving memory and compute on long sequences. However, they approximate the full attention matrix and can miss some token interactions, potentially reducing accuracy or requiring more training to recover quality. This trade-off is the key consideration.
- ✗
It eliminates the need for positional encodings because approximate attention is order-invariant.
Why it's wrong here
Approximate attention methods still require positional information to distinguish token order. Without positional encodings, the model cannot tell whether tokens appear early or late in the sequence. These methods approximate the attention computation, not the need for order awareness. Removing positional encodings would severely degrade performance on language tasks.
- ✗
It increases the number of parameters in the model, which always improves accuracy.
Why it's wrong here
Approximate attention typically changes the computation pattern, not the parameter count. Adding parameters is a separate scaling decision and does not guarantee better accuracy, especially if the approximation loses information. The goal is efficiency, and more parameters would counteract that by increasing memory and compute costs.
- ✗
It makes the model immune to the vanishing gradient problem in deep transformer stacks.
Why it's wrong here
Vanishing gradients in deep transformers are addressed by residual connections, layer normalization, and careful initialization, not by approximate attention. Approximate attention targets computational complexity for long sequences. It does not inherently solve gradient flow issues, and claiming immunity is technically incorrect.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.