NCP-GENL LLM Architecture Practice Question
An engineer must serve a 70B-parameter LLM for a workload with many concurrent users and long shared system prompts, and wants to maximize throughput without retraining. Which inference-time optimization most directly reduces redundant computation across requests sharing the same prompt prefix?
⚠ Common exam trap
The trap here is conflating memory-footprint optimizations such as KV cache quantization with compute-deduplication techniques like prefix caching, since both involve the KV cache but solve different bottlenecks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Prefix caching of key/value tensors for shared prompt prefixes
When many requests share a long system prompt, recomputing attention over that prefix for every request wastes prefill compute. Prefix caching persists the key/value tensors for the shared prefix and reuses them, cutting redundant work and improving throughput under concurrency. Speculative decoding targets decode latency, beam search adds work, and KV cache quantization saves memory but not duplicated prefill computation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increasing the beam width during generation
Why it's wrong here
Beam search explores multiple candidate continuations simultaneously, which multiplies compute and memory per request. It improves output quality in some tasks but increases total work rather than removing duplicated prefix processing, so it works against the throughput goal in this high-concurrency scenario.
- ✗
Speculative decoding with a smaller draft model
Why it's wrong here
Speculative decoding accelerates generation by proposing multiple tokens with a draft model and verifying them in parallel with the target model. It reduces per-token latency for decoding but does not deduplicate the prefill computation of an identical prompt prefix across many concurrent requests.
- ✗
Enabling FP8 quantization of the KV cache
Why it's wrong here
FP8 KV cache quantization reduces the memory footprint of cached keys and values, allowing more concurrent sequences to fit in VRAM. It improves capacity and can indirectly help throughput, but it does not eliminate the redundant computation of identical prompt prefixes, which is the specific inefficiency described.
- ✓
Prefix caching of key/value tensors for shared prompt prefixes
Why this is correct
Prefix caching stores the key/value tensors computed for a shared prefix, such as a long system prompt, and reuses them across requests instead of recomputing attention for those tokens on every call. This directly eliminates the redundant prefill work the scenario describes and is supported in NVIDIA TensorRT-LLM and similar serving stacks.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.