NCP-GENL Model Deployment Practice Question
A team is deploying a 13B-parameter LLM with NVIDIA TensorRT-LLM on a single A100 80GB GPU. They want to reduce GPU memory usage during inference without retraining the model, while keeping acceptable output quality. Which technique should they apply?
⚠ Common exam trap
The trap here is assuming that KV cache tuning or sequence-length reduction will solve weight-dominated memory pressure for a large model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable INT8 weight-only quantization using TensorRT-LLM's quantization toolkit.
INT8 weight-only quantization is a post-training technique supported by TensorRT-LLM that reduces weight memory by roughly half while preserving output quality. It directly addresses the goal of lowering GPU memory usage for a large model on a single GPU without retraining. Other listed changes affect scheduling or activation memory rather than the core weight footprint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable INT8 weight-only quantization using TensorRT-LLM's quantization toolkit.
Why this is correct
INT8 weight-only quantization compresses model weights to 8-bit integers while keeping activations in higher precision, cutting weight memory roughly in half with minimal quality loss. TensorRT-LLM supports this via its quantization toolkit and calibration workflow, making it a practical post-training approach for a 13B model on a single 80GB GPU without retraining.
- ✗
Reduce the maximum sequence length to 512 tokens to lower activation memory.
Why it's wrong here
Lowering the maximum sequence length reduces activation and KV cache memory somewhat, but the dominant memory consumer for a 13B model is the weight storage. This change constrains functionality and does not substantially reduce the model's parameter memory. It is a partial mitigation, not the targeted technique requested.
- ✗
Increase the KV cache block size to 128 tokens to reduce memory fragmentation.
Why it's wrong here
Larger KV cache block sizes can reduce fragmentation overhead but do not shrink the model weights, which dominate memory for a 13B model. This change affects paged attention memory management, not the core memory footprint of the model parameters. It would not meaningfully address the goal of reducing overall GPU memory usage.
- ✗
Enable Tensor Parallelism across two GPUs to split the model.
Why it's wrong here
Tensor Parallelism splits a model across multiple GPUs but the scenario specifies a single A100 80GB GPU. Adding GPUs changes the hardware requirement rather than reducing memory usage on the existing device. The question asks for a technique that works without additional hardware, so this does not satisfy the constraint.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.