NCP-GENL LLM Architecture • 10 Questions
10 NCP-GENL LLM Architecture practice questions with answers and explanations. Free, no signup.
An engineer is analyzing a 70B-parameter decoder-only LLM that uses Grouped Query Attention (GQA) with 8 key/value heads and 64 query heads. During inference with a long prompt, they notice that the KV cache memory is still a bottleneck on a single GPU, and they want to reduce it further without retraining the model from scratch. Which approach best describes a valid architectural modification that reduces KV cache size while preserving the model's learned attention behavior as much as possible?
Choose an answer to begin — your selection is scored in the full session.
10 questions · instant feedback and full explanations after every question.