Courseiva
LLM Architecture →mediumMultiple Choice

NCP-GENL LLM Architecture Practice Question

An engineer is deploying a 70B parameter model and needs to serve many concurrent users on a single GPU with limited memory. They want to store the attention keys and values for past tokens efficiently so that generation does not recompute them at every step. Which technique should they implement?

⚠ Common exam trap

The trap here is treating training-time memory tricks like gradient checkpointing as if they accelerated autoregressive inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A key-value cache that stores the projected keys and values for previously processed tokens.

Autoregressive decoding generates one token at a time, and without caching the model would recompute keys and values for the entire prefix at every step, wasting compute. A key-value cache stores those projections so each step only processes the new token, which is the standard way to make long-context generation practical.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Quantization-aware training that reduces weight precision before deployment.

    Why it's wrong here

    Quantization-aware training reduces the memory footprint of weights and can speed up matrix multiplications, but it does not change the fact that past keys and values must be stored or recomputed. It is orthogonal to the caching requirement in this scenario.

  • ✓

    A key-value cache that stores the projected keys and values for previously processed tokens.

    Why this is correct

    A key-value cache retains the key and value projections for all past tokens so each new step only computes the query and attends to cached entries. This avoids recomputing the full prefix at every decoding step, dramatically reducing compute for autoregressive generation.

  • ✗

    Speculative decoding that uses a smaller draft model to propose tokens.

    Why it's wrong here

    Speculative decoding can improve throughput by verifying multiple candidate tokens, but it does not eliminate repeated computation of past keys and values. Without a key-value cache, each verification step still recomputes the prefix, so memory and compute remain high.

  • ✗

    Gradient checkpointing that recomputes activations during the backward pass.

    Why it's wrong here

    Gradient checkpointing is a training-time memory optimization that trades compute for activation storage. It does not help inference latency or avoid recomputation of past keys and values during autoregressive decoding, so it does not address the scenario.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.