NCP-GENL LLM Architecture Practice Question
An inference engineer is serving a 70B-parameter decoder-only LLM and wants to reduce KV cache memory to fit longer contexts on each GPU. They consider Multi-Query Attention (MQA), Grouped-Query Attention (GQA), and standard Multi-Head Attention (MHA). Which statement correctly describes the memory and quality tradeoff among these attention variants?
⚠ Common exam trap
The trap here is conflating GQA's grouped key-value heads with MQA's single shared head, which understates GQA's memory footprint.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
MQA shares key and value projections across all query heads, giving the smallest KV cache but often degrading quality relative to MHA.
MQA minimizes KV cache by sharing one K and V head across all query heads, at some cost to quality. GQA interpolates by grouping query heads with dedicated K and V heads, balancing memory and quality. MHA provides the richest representation but the largest cache. The other options misstate GQA's structure, reverse MHA's memory ranking, or incorrectly claim GQA removes the cache.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
GQA eliminates the KV cache entirely by recomputing keys and values on the fly during decoding.
Why it's wrong here
GQA still caches keys and values; it simply groups query heads so fewer distinct K and V heads need storage. Recomputing keys and values on every decoding step would negate the purpose of the cache and dramatically increase compute. No mainstream attention variant eliminates the KV cache while retaining efficient autoregressive decoding, so this option misrepresents how GQA works.
- ✗
GQA uses a single key and value head shared by all query heads, matching MQA memory savings while preserving MHA quality.
Why it's wrong here
GQA partitions query heads into groups, each with its own key and value head, so it sits between MHA and MQA. Describing it as a single shared head conflates it with MQA. GQA reduces KV cache relative to MHA but does not match MQA's minimal footprint, and its quality is generally closer to MHA, not identical, so the claim is inaccurate on both counts.
- ✗
MHA has the smallest KV cache because each head stores its own keys and values independently.
Why it's wrong here
Storing keys and values independently per head is exactly what makes MHA's KV cache the largest among the three variants. The number of K and V heads equals the number of query heads, multiplying cache size. MQA and GQA were introduced specifically to shrink this footprint, so claiming MHA is smallest reverses the actual memory ordering.
- ✓
MQA shares key and value projections across all query heads, giving the smallest KV cache but often degrading quality relative to MHA.
Why this is correct
In MQA, every query head uses the same key and value head, so the KV cache stores only one K and one V vector per layer per token, dramatically reducing memory. This aggressive sharing can hurt model quality and training stability compared with MHA, which is why intermediate GQA designs were introduced. The statement accurately captures both the memory benefit and the quality risk.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.