Courseiva
LLM Architecture →hardMultiple Choice

NCP-GENL LLM Architecture Practice Question

An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?

⚠ Common exam trap

The trap here is conflating memory-footprint optimizations such as KV cache quantization with compute-deduplication techniques like prefix caching, since both involve the KV cache but solve different bottlenecks.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Prefix caching of key/value tensors for shared prompt prefixes

When many requests share a long system prompt, recomputing attention over that prefix for every request wastes prefill compute. Prefix caching persists the key/value tensors for the shared prefix and reuses them, cutting redundant work and improving throughput under concurrency. Speculative decoding targets decode latency, beam search adds work, and KV cache quantization saves memory but not duplicated prefill computation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increasing the beam width during generation

    Why it's wrong here

    Beam search explores multiple candidate continuations simultaneously, which multiplies compute and memory per request. It improves output quality in some tasks but increases total work rather than removing duplicated prefix processing, so it works against the throughput goal in this high-concurrency scenario.

  • ✗

    Speculative decoding with a smaller draft model

    Why it's wrong here

    Speculative decoding accelerates generation by proposing multiple tokens with a draft model and verifying them in parallel with the target model. It reduces per-token latency for decoding but does not deduplicate the prefill computation of an identical prompt prefix across many concurrent requests.

  • ✗

    Enabling FP8 quantization of the KV cache

    Why it's wrong here

    FP8 KV cache quantization reduces the memory footprint of cached keys and values, allowing more concurrent sequences to fit in VRAM. It improves capacity and can indirectly help throughput, but it does not eliminate the redundant computation of identical prompt prefixes, which is the specific inefficiency described.

  • ✓

    Prefix caching of key/value tensors for shared prompt prefixes

    Why this is correct

    Prefix caching stores the key/value tensors computed for a shared prefix, such as a long system prompt, and reuses them across requests instead of recomputing attention for those tokens on every call. This directly eliminates the redundant prefill work the scenario describes and is supported in NVIDIA TensorRT-LLM and similar serving stacks.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.