NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A machine learning engineer is deploying a transformer-based language model for real-time translation. They observe that inference latency is too high for the required throughput. The model uses standard multi-head self-attention. Which modification is most likely to reduce latency without significantly degrading translation quality?
⚠ Common exam trap
The trap here is thinking more capacity (heads, dimensions, layers) improves performance without considering the latency penalty, when the goal is to reduce latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a sparse attention mechanism
Sparse attention reduces the quadratic cost of self-attention by limiting each token's attention to a subset of positions, which lowers computation and latency. Increasing heads, hidden dimension, or layers all raise computational load and latency. Sparse attention is a targeted optimization for real-time inference while preserving translation quality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of attention heads
Why it's wrong here
Adding more attention heads increases the computational cost of the self-attention mechanism, as each head performs its own scaled dot-product attention. This would raise latency rather than reduce it, and may not improve quality if the model is already well-tuned. For real-time translation, it is counterproductive to the goal of lower latency.
- ✓
Use a sparse attention mechanism
Why this is correct
Sparse attention reduces the number of key-value pairs each query attends to, lowering the quadratic complexity of self-attention to near-linear. This directly cuts computation and memory access, decreasing latency. For translation, where local context often suffices, sparse patterns can maintain quality while meeting real-time throughput requirements.
- ✗
Increase the model's hidden dimension
Why it's wrong here
Expanding the hidden dimension increases the size of weight matrices and the cost of matrix multiplications, leading to higher latency and memory usage. It may improve capacity but does not address the real-time constraint. In this scenario, it would worsen the problem rather than solve it.
- ✗
Add more layers to the transformer
Why it's wrong here
Adding more layers increases the depth of the network, requiring more sequential computation and thus higher latency. While deeper models can capture more complex patterns, the trade-off is slower inference. For real-time translation, this is not a viable way to reduce latency and may also increase memory footprint.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.