Courseiva
Model Deployment →mediumMultiple Select

NCP-GENL Model Deployment Practice Question

A team is deploying a large language model on NVIDIA Triton Inference Server with TensorRT-LLM backend. They want to reduce GPU memory consumption to fit a larger model on a single GPU without significantly degrading output quality. Which two techniques should they use? (Choose two.)

⚠ Common exam trap

The trap here is assuming that increasing batch size or using FP32 improves memory efficiency, when they actually increase memory usage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable INT8 weight-only quantization

INT8 weight-only quantization reduces weight memory footprint with minimal quality loss, and paged KV cache optimizes memory usage during generation by reducing fragmentation. Together, they enable larger models to fit on a single GPU. FP32 increases memory, larger batch sizes demand more memory, and tensor parallelism requires multiple GPUs, so they do not meet the goal.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the maximum batch size

    Why it's wrong here

    Increasing the maximum batch size allows more concurrent requests but increases memory usage because each request requires its own KV cache and activations. This would exacerbate memory constraints rather than alleviate them. It is not a technique to reduce memory consumption; instead, it demands more memory.

  • ✗

    Use FP32 precision for all computations

    Why it's wrong here

    FP32 precision doubles memory usage compared to FP16 and does not reduce memory consumption. It is typically used for training or when numerical stability is critical, but for inference, it increases memory footprint and reduces throughput. It would not help fit a larger model on a single GPU.

  • ✗

    Use tensor parallelism across multiple GPUs

    Why it's wrong here

    Tensor parallelism splits the model across multiple GPUs, which reduces per-GPU memory usage but requires multiple GPUs. The scenario specifies fitting a larger model on a single GPU, so this technique is not applicable. It also introduces communication overhead and does not address single-GPU memory reduction.

  • ✓

    Enable INT8 weight-only quantization

    Why this is correct

    INT8 weight-only quantization reduces the precision of model weights from FP16 to INT8, cutting memory usage by roughly half. It often preserves output quality well because activations remain in higher precision. This allows larger models to fit on a single GPU with minimal accuracy loss, making it a suitable technique for the scenario.

  • ✓

    Enable paged KV cache

    Why this is correct

    Paged KV cache stores key/value tensors in non-contiguous memory blocks, reducing fragmentation and allowing more efficient memory utilization. It can significantly lower memory overhead during generation, enabling larger batch sizes or models. It does not degrade output quality and is a standard optimization in TensorRT-LLM.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.