An engineer is analyzing a 70B-parameter decoder-only LLM that uses Grouped Query Attention (GQA) with 8 key/value heads and 64 query heads. During inference with a long prompt, they notice that the KV cache memory is still a bottleneck on a single GPU, and they want to reduce it further without retraining the model from scratch. Which approach best describes a valid architectural modification that reduces KV cache size while preserving the model's learned attention behavior as much as possible?
Trap 1: Replace GQA with Multi-Query Attention (MQA) by using a single…
MQA uses one key/value head, which drastically reduces KV cache size, but switching from GQA to MQA without training changes the attention computation and typically degrades quality. The model was trained with 8 key/value heads, so collapsing them to one without adaptation breaks learned behavior. This violates the requirement to preserve learned attention behavior and avoid retraining from scratch.
Trap 2: Reduce the model's hidden dimension by pruning entire attention…
Pruning entire attention heads changes the model's capacity and requires fine-tuning to recover quality, which is a form of retraining. It also reduces the number of query heads and can degrade the model's learned representations. While it may reduce KV cache size, it does not preserve learned attention behavior as well as lower-precision storage and is not a post-training-only modification.
Trap 3: Increase the number of key/value heads to match the number of query…
Increasing key/value heads to match query heads converts the model to full multi-head attention, which increases KV cache size proportionally. This directly worsens the memory bottleneck rather than reducing it. It also changes the attention computation and would require retraining or fine-tuning to preserve quality, contradicting the goal of reducing cache without retraining from scratch.
- A
Replace GQA with Multi-Query Attention (MQA) by using a single key/value head for all query heads, without any additional training.
Why it fails: MQA uses one key/value head, which drastically reduces KV cache size, but switching from GQA to MQA without training changes the attention computation and typically degrades quality. The model was trained with 8 key/value heads, so collapsing them to one without adaptation breaks learned behavior. This violates the requirement to preserve learned attention behavior and avoid retraining from scratch.
- B
Apply weight-only quantization to the key/value projection weights and store the KV cache in a lower-precision format such as FP8 or INT8.
Storing the KV cache in FP8 or INT8 reduces its memory footprint by roughly 2x to 4x compared to FP16, directly alleviating the bottleneck. Weight-only quantization of the key/value projections can be applied post-training with minimal impact on attention behavior. This approach preserves the GQA architecture and the number of heads, so it does not require retraining the model from scratch.
- C
Reduce the model's hidden dimension by pruning entire attention heads across all layers and fine-tuning only the remaining heads.
Why it fails: Pruning entire attention heads changes the model's capacity and requires fine-tuning to recover quality, which is a form of retraining. It also reduces the number of query heads and can degrade the model's learned representations. While it may reduce KV cache size, it does not preserve learned attention behavior as well as lower-precision storage and is not a post-training-only modification.
- D
Increase the number of key/value heads to match the number of query heads, converting GQA to multi-head attention.
Why it fails: Increasing key/value heads to match query heads converts the model to full multi-head attention, which increases KV cache size proportionally. This directly worsens the memory bottleneck rather than reducing it. It also changes the attention computation and would require retraining or fine-tuning to preserve quality, contradicting the goal of reducing cache without retraining from scratch.