NCP-GENL Model Deployment Practice Question
A team is deploying a large language model on NVIDIA Triton Inference Server with NVIDIA TensorRT-LLM backend. They need to reduce GPU memory usage to fit a larger model on the same hardware while maintaining acceptable latency. Which two techniques should they use? (Choose two.)
⚠ Common exam trap
The trap here is thinking that increasing tensor parallel size reduces memory on the same hardware, when it actually requires additional GPUs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Quantize the model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.
Quantizing weights to INT8 or FP8 reduces memory footprint significantly, and paged KV cache with in-flight batching optimizes runtime memory usage. Together they allow larger models or longer sequences on the same GPU. Other options either require more hardware, increase memory usage, or affect scheduling rather than memory footprint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the Triton dynamic batching max_queue_delay_microseconds to 0 to process requests immediately.
Why it's wrong here
Setting max_queue_delay_microseconds to 0 disables batching benefits and can increase the number of concurrent sequences waiting, potentially increasing memory pressure. It does not reduce memory usage. Dynamic batching parameters control latency and throughput, not memory footprint. This would not help fit a larger model and may worsen memory utilization.
- ✓
Quantize the model weights to INT8 or FP8 using TensorRT-LLM's quantization toolkit.
Why this is correct
Quantizing weights to INT8 or FP8 reduces the memory footprint by 2x to 4x compared to FP16, allowing larger models to fit. TensorRT-LLM supports post-training quantization and quantization-aware training. This directly addresses the memory constraint while maintaining acceptable latency, as quantized kernels are optimized for NVIDIA GPUs. It is a standard technique for memory-constrained deployments.
- ✗
Increase the tensor parallel size to 8 across eight GPUs.
Why it's wrong here
Increasing tensor parallel size spreads the model across more GPUs, which reduces per-GPU memory but requires more hardware. The scenario specifies fitting a larger model on the same hardware, so adding GPUs is not an option. It also increases communication overhead. This does not solve the memory constraint without additional resources.
- ✓
Enable paged KV cache and in-flight batching in the TensorRT-LLM backend.
Why this is correct
Paged KV cache reduces memory fragmentation by allocating KV cache in fixed-size blocks, and in-flight batching allows dynamic reuse of those blocks as sequences finish. Together they improve memory utilization and enable more concurrent sequences without increasing total memory. This is a runtime optimization that complements weight quantization. It helps fit larger models or longer sequences on the same GPU.
- ✗
Use FP32 precision for all layers to avoid quantization errors.
Why it's wrong here
FP32 precision doubles memory usage compared to FP16 and quadruples compared to INT8. It is the opposite of what is needed to reduce memory. While it may improve numerical accuracy, it directly conflicts with the goal of fitting a larger model on the same hardware. TensorRT-LLM deployments typically use FP16 or lower precision for efficiency.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.