Courseiva
LLM Architecture →mediumMultiple Choice

NCP-GENL LLM Architecture Practice Question

A team is deploying a 70B-parameter decoder-only LLM on an NVIDIA H100 GPU node. During generation they observe that the KV cache grows linearly with sequence length and is consuming most of the available HBM, forcing them to limit batch size. They want to reduce KV cache memory without retraining the model from scratch. Which architectural change should they apply?

⚠ Common exam trap

The trap here is assuming that any attention efficiency change such as RoPE or more heads reduces KV cache size, when only reducing the number of KV heads actually shrinks the cache.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Replace multi-head attention with Grouped Query Attention (GQA), where multiple query heads share a single key/value head.

Grouped Query Attention reduces the number of key/value heads relative to query heads, so the KV cache that must be stored for each token shrinks by the grouping factor. Because the cache dominates HBM during long-context decoding, this architectural change directly relieves memory pressure and permits larger batches, and it can be adopted via uptraining rather than a full retrain.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Switch the feed-forward network activation from SwiGLU to ReLU to reduce parameter count.

    Why it's wrong here

    The feed-forward network activation choice affects parameter count and compute in the MLP blocks, not the size of the key/value tensors stored during autoregressive decoding. The KV cache is entirely a function of attention projections, so changing the MLP activation does nothing to reduce the memory consumed by cached keys and values across the sequence.

  • ✗

    Apply rotary positional embeddings to the key and query vectors instead of learned absolute embeddings.

    Why it's wrong here

    Rotary positional embeddings change how position information is injected into queries and keys, improving length generalization and relative position encoding, but they do not reduce the number of key/value tensors that must be cached. The cache size is determined by the count and dimension of KV heads, not by the positional encoding scheme used on them.

  • ✗

    Increase the number of attention heads while keeping the head dimension constant.

    Why it's wrong here

    Adding more attention heads while holding head dimension constant increases the total hidden size and, more importantly, increases the number of key and value projections. The KV cache scales with the number of KV heads times head dimension times layers, so this change would enlarge the cache rather than reduce it, worsening the HBM pressure the team is trying to solve.

  • ✓

    Replace multi-head attention with Grouped Query Attention (GQA), where multiple query heads share a single key/value head.

    Why this is correct

    GQA reduces the number of distinct key and value projections by having groups of query heads share KV heads, which directly shrinks the KV cache proportionally to the number of KV heads rather than query heads. This lowers HBM pressure during inference and allows larger batch sizes, and it can be introduced through continued pretraining or uptraining rather than a full retrain from scratch.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.