NCP-GENL Model Deployment Practice Question
A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)
⚠ Common exam trap
The trap here is assuming that increasing batch size or using FP32 improves memory efficiency, when they actually increase memory usage.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable INT8 weight-only quantization
INT8 weight-only quantization reduces weight memory footprint with minimal quality loss, and paged KV cache optimizes memory usage during generation by reducing fragmentation. Together, they enable larger models to fit on a single GPU. FP32 increases memory, larger batch sizes demand more memory, and tensor parallelism requires multiple GPUs, so they do not meet the goal.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the maximum batch size
Why it's wrong here
Increasing the maximum batch size allows more concurrent requests but increases memory usage because each request requires its own KV cache and activations. This would exacerbate memory constraints rather than alleviate them. It is not a technique to reduce memory consumption; instead, it demands more memory.
- ✗
Use FP32 precision for all computations
Why it's wrong here
FP32 precision doubles memory usage compared to FP16 and does not reduce memory consumption. It is typically used for training or when numerical stability is critical, but for inference, it increases memory footprint and reduces throughput. It would not help fit a larger model on a single GPU.
- ✗
Use tensor parallelism across multiple GPUs
Why it's wrong here
Tensor parallelism splits the model across multiple GPUs, which reduces per-GPU memory usage but requires multiple GPUs. The scenario specifies fitting a larger model on a single GPU, so this technique is not applicable. It also introduces communication overhead and does not address single-GPU memory reduction.
- ✓
Enable INT8 weight-only quantization
Why this is correct
INT8 weight-only quantization reduces the precision of model weights from FP16 to INT8, cutting memory usage by roughly half. It often preserves output quality well because activations remain in higher precision. This allows larger models to fit on a single GPU with minimal accuracy loss, making it a suitable technique for the scenario.
- ✓
Enable paged KV cache
Why this is correct
Paged KV cache stores key/value tensors in non-contiguous memory blocks, reducing fragmentation and allowing more efficient memory utilization. It can significantly lower memory overhead during generation, enabling larger batch sizes or models. It does not degrade output quality and is a standard optimization in TensorRT-LLM.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.